arXiv AI By Yuetong Wu, Maojun Sun

AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis

Read the original on arXiv AI →

AutoXRD is an autonomous large language model (LLM) agent framework designed to automate powder X-ray diffraction (XRD) analysis by structuring the process as stepwise refinement, grounding actions in observed evidence, and applying deterministic crystallographic and physical checks. The authors introduce XRDBench, comprising two tracks: XRDBench-QA with 100 diagnostic tasks focused on scientific reasoning, and XRDBench-E2E with 34 executable workflows that test full analysis capabilities, including file inspection, software execution, iterative refinement, evidence preservation, and reporting. Evaluation of ten recent LLMs on 1,340 model–task runs shows average scores of 57.8, with GPT‑5.6 Sol achieving the highest overall score of 81.1; the study also identifies key failure modes such as coupled‑parameter control and quantitative reasoning, highlighting areas for future improvement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

CompMat-Bench is a new benchmark comprising 94 tasks drawn from recent computational materials science studies, designed to evaluate AI agents on realistic research steps without requiring costly simulations during testing. The benchmark pre‑reproduces inputs and outputs to provide ground truth, allowing agents to be graded with fixed rules rather than an LLM judge. It supports both single tasks and multi‑step workflows, with varying levels of methodological guidance, and shows that while agents can achieve high pass rates on individual tasks, performance drops in longer workflows or with reduced guidance, often due to scientific rather than software errors.

By Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson
arXiv AI
Jun 6

AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations

arXiv:2605. 26179v2 Announce Type: replace-cross Abstract: Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting algorithms when convergence stalls, revising plans when unexpected physics emerges, and inserting steps as intermediate results reshape the problem.

By Penghui Yang, Zhonghan Zhang, Yue Li, Xinrun Wang, Yanchen Deng, Yuhao Lu, Bijun Tang, Zheng Liu, Bo An