arXiv AI

Discovering physical mechanisms from experiment-simulation mismatches

The paper introduces eXplainable DFT (XDFT), a self‑evolving computational agent that transforms experiment‑simulation mismatches into executable searches for physical mechanisms. XDFT formalizes candidate mechanisms as hypotheses, tests them against experimental data, and refines its search strategy through a learning loop. In a benchmark of 112 cases where standard calculations predicted a metal but experiments found a semiconductor, XDFT resolved 105 cases with evidence‑supported mechanisms, and its top‑ranked hypotheses improved dramatically over initial expert priors.

arXiv AI
Jun 6

AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations

arXiv:2605. 26179v2 Announce Type: replace-cross Abstract: Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting algorithms when convergence stalls, revising plans when unexpected physics emerges, and inserting steps as intermediate results reshape the problem.

By Penghui Yang, Zhonghan Zhang, Yue Li, Xinrun Wang, Yanchen Deng, Yuhao Lu, Bijun Tang, Zheng Liu, Bo An
arXiv AI
Sep 24

False-science induction in autonomous scientific discovery

The paper investigates how closed‑loop autonomous discovery systems can develop false‑science induction when physical objects and measurements are incorrectly paired. It demonstrates that such misbinding causes neural surrogates to learn spurious associations, diverting experimental effort toward low‑performing regions in both green fluorescent protein fitness and materials band‑gap prediction loops. The study shows that the coherence of these errors—not just their frequency—drives budget misallocation and proposes monitoring strategies to detect and quarantine corrupted hypothesis axes.

By Hanbing Liang, Fujun Liu
arXiv AI
Sep 12

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

The paper introduces ARCHE, an autonomous system that combines a general-purpose reasoning model, a domain-specialized computational chemistry model, and a structured tool registry to automate chemical mechanism discovery. ARCHE interprets scientific questions, generates and prioritizes mechanistic hypotheses, orchestrates computational workflows, and refines conclusions in a closed loop. The authors validate the system on three challenging scenarios, including reconstructing stereocontrolling transition states, proposing a radical pathway for an unpublished reaction, and identifying a descriptor governing selectivity in nickel-catalyzed cross‑coupling reactions.

By Dong Li, Sixuan Mi, Zihao Ye, Huan Xiong, Tao XU, Tong Zhu, Aijia Zhang, Junqi Gao, Kaiyan Zhang, Shijie Wang, Bowen Zhou, Yuqiang Li, Biqing Qi
arXiv AI
Jun 30

Hierarchical Experimentalist Agents

arXiv:2606. 29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search.

By Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka, Varun Gandhi, Scott Niekum
arXiv AI
Jun 18

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

arXiv:2606. 18874v1 Announce Type: new Abstract: AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference.

By Zijian Wang, Hanqi Li, Ziyue Yang, Zijian Hu, Shenghan Zuo, Yunzhe Zhang, Da Ma, Danyu Luo, Chenrun Wang, Jing Peng, Tiancheng Huang, Sijia Guo, Huayang Wang, Zichen Zhu, Senyu Han, Yilu Cao, Kai Yu, Lu Chen
arXiv AI
Aug 28

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

The paper "Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research" argues that large language model agents must faithfully implement reference methods, design experiments that truly test claims, and provide supporting evidence. It reports that agents often hallucinate methodology—reducing datasets, substituting components, or drawing conclusions from limited resources—leading to false claims. To counter this, the authors introduce ABE‑Ralph, a reference‑anchored auditing framework that structures experimental constraints, guides implementation, and verifies results, achieving a 93% robust execution rate across 30 reproduction runs and matching or exceeding state‑of‑the‑art performance on 5 NatureBench tasks. "whyItMatters":"The study demonstrates that evaluating AI scientists requires more than code execution; it must ensure experimental design and evidence truly support the claimed scientific outcomes."

By Lezhi Yu, Xiaogang Xu, Yuhua Zhou, Shuibing He, Aimin Pan