arXiv AI

AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org

arXiv:2512. 11935v2 Announce Type: replace Abstract: Agentic AI systems increasingly connect large language models (LLMs) to external scientific tools, yet whether and when tool access improves prediction accuracy remains uncharacterized.

arXiv AI
2d ago

CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

CompMat-Bench is a new benchmark comprising 94 tasks drawn from recent computational materials science studies, designed to evaluate AI agents on realistic research steps without requiring costly simulations during testing. The benchmark pre‑reproduces inputs and outputs to provide ground truth, allowing agents to be graded with fixed rules rather than an LLM judge. It supports both single tasks and multi‑step workflows, with varying levels of methodological guidance, and shows that while agents can achieve high pass rates on individual tasks, performance drops in longer workflows or with reduced guidance, often due to scientific rather than software errors.

By Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson
arXiv AI
Jun 6

AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations

arXiv:2605. 26179v2 Announce Type: replace-cross Abstract: Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting algorithms when convergence stalls, revising plans when unexpected physics emerges, and inserting steps as intermediate results reshape the problem.

By Penghui Yang, Zhonghan Zhang, Yue Li, Xinrun Wang, Yanchen Deng, Yuhao Lu, Bijun Tang, Zheng Liu, Bo An
arXiv AI
Sep 15

El Agente Potente: High-Throughput Agentic Atomistic Simulations

El Agente Potente is an agentic system that integrates typed execution graphs and a coding mode to facilitate machine‑learning interatomic potential (MLIP) driven atomistic simulations. Typed execution graphs offer structured, provenance‑aware workflows where large language models handle planning and routing while deterministic Python code performs scientific computation and validation. The coding agent builds customized workflows for tasks needing procedural flexibility, invoking existing Potente functions for supported calculations. The system is demonstrated across materials discovery, energy‑landscape exploration, adsorption, and catalytic reaction workflows, with benchmarks on reproducibility and LLM token cost.

By Tsz Wai Ko, Jiaru Bai, Thomas Swanick, Yeonghun Kang, Changhyeok Choi, Angelina Qihong Jiang, Aiwei Yin, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv AI
Aug 13

DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation

arXiv:2507. 14267v2 Announce Type: replace Abstract: Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents lose context, game verification checks, and can produce large volumes of plausible yet invalid results.

By Ziqi Wang, Hongshuo Huang, Hancheng Zhao, Changwen Xu, Shang Zhu, Jan Janssen, Venkatasubramanian Viswanathan
arXiv AI
Jul 28

An Agentic Orchestration of Atomistic Simulations

arXiv:2607. 22596v1 Announce Type: new Abstract: Atomistic simulations are central to materials design, but their execution involves complex, multi-step workflows that require significant human expertise.

By Rahul Somasundaram, Adela Habib, Khanh Dang, Sachin Shivakumar, Ryley G. Hill, Golo Wimmer, Avanish Mishra, Aleksandra Pachalieva, Arthur Lui, Hari Viswanathan, Michael Grosskopf, Saryu Fensin, Russell Bent, Nathan DeBardeleben, Earl Lawrence
Hugging Face Trending Papers
Aug 4

Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design.

arXiv AI
Sep 24

MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

MolDesignBench is a new benchmark for evaluating large language model (LLM)-based agents in scenario‑grounded molecular design. It contains 2,000 generation and optimization tasks that blend implicit narrative requirements with explicit property and functional‑group constraints, including infeasible cases, and require the use of 17 specialized chemistry tools. Experiments with leading LLMs show low success rates (best ~43%) and highlight failures in implicit‑constraint reasoning, infeasibility detection, and tool usage, underscoring the benchmark’s role in identifying key bottlenecks for future research.

By Yongjun Jeong, Hanbum Ko, Ye Rin Kim, Chanhui Lee, Rodrigo Hormazabal, Jaewan Lee, Sehui Han, Sungbin Lim, Sungwoong Kim
arXiv AI
Aug 18

ALKEMIE Agent: an autonomous platform for computational materials design

arXiv:2608. 15776v1 Announce Type: cross Abstract: Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material computation workflows remain fragmented and heavily manual, requiring researchers to constantly bridge software tools, data analysis, and intermediate decisions.

By Hongfu Huang, Yuzhe Li, Ao Xu, Bo Liu, Changrui Wang, Kan Tang, Ning Yang, Shengxian Liu, Hanyu Liu, Pengpeng Zhang, Linggang Zhu, Fengkai Liu, Yichen Lu, Tong Zhao, Naihua Miao, Jian Zhou, Zhimei Sun
arXiv AI
Sep 2

AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis

AutoXRD is an autonomous large language model (LLM) agent framework designed to automate powder X-ray diffraction (XRD) analysis by structuring the process as stepwise refinement, grounding actions in observed evidence, and applying deterministic crystallographic and physical checks. The authors introduce XRDBench, comprising two tracks: XRDBench-QA with 100 diagnostic tasks focused on scientific reasoning, and XRDBench-E2E with 34 executable workflows that test full analysis capabilities, including file inspection, software execution, iterative refinement, evidence preservation, and reporting. Evaluation of ten recent LLMs on 1,340 model–task runs shows average scores of 57.8, with GPT‑5.6 Sol achieving the highest overall score of 81.1; the study also identifies key failure modes such as coupled‑parameter control and quantitative reasoning, highlighting areas for future improvement.

By Yuetong Wu, Maojun Sun