The paper introduces eXplainable DFT (XDFT), a self‑evolving computational agent that transforms experiment‑simulation mismatches into executable searches for physical mechanisms. XDFT formalizes candidate mechanisms as hypotheses, tests them against experimental data, and refines its search strategy through a learning loop. In a benchmark of 112 cases where standard calculations predicted a metal but experiments found a semiconductor, XDFT resolved 105 cases with evidence‑supported mechanisms, and its top‑ranked hypotheses improved dramatically over initial expert priors.
By Yue Li, Penghui Yang, Yushan Xiao, Zhonghan Zhang, Jianguo Huang, Yuhao Lu, Cuntai Guan, Bo An, Bijun Tang, Zheng Liu
arXiv:2605. 10246v2 Announce Type: replace Abstract: AI scientist systems are increasingly deployed for autonomous research, yet their academic integrity has never been systematically evaluated.
By Zonglin Yang, Xingtong Liu, Xinyan Xu
arXiv:2606. 18874v1 Announce Type: new Abstract: AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference.
By Zijian Wang, Hanqi Li, Ziyue Yang, Zijian Hu, Shenghan Zuo, Yunzhe Zhang, Da Ma, Danyu Luo, Chenrun Wang, Jing Peng, Tiancheng Huang, Sijia Guo, Huayang Wang, Zichen Zhu, Senyu Han, Yilu Cao, Kai Yu, Lu Chen
The paper "Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research" argues that large language model agents must faithfully implement reference methods, design experiments that truly test claims, and provide supporting evidence. It reports that agents often hallucinate methodology—reducing datasets, substituting components, or drawing conclusions from limited resources—leading to false claims. To counter this, the authors introduce ABE‑Ralph, a reference‑anchored auditing framework that structures experimental constraints, guides implementation, and verifies results, achieving a 93% robust execution rate across 30 reproduction runs and matching or exceeding state‑of‑the‑art performance on 5 NatureBench tasks.
"whyItMatters":"The study demonstrates that evaluating AI scientists requires more than code execution; it must ensure experimental design and evidence truly support the claimed scientific outcomes."
By Lezhi Yu, Xiaogang Xu, Yuhua Zhou, Shuibing He, Aimin Pan
arXiv:2606. 11851v1 Announce Type: new Abstract: Open-ended scientific discovery asks agents to move beyond executing analyses for predefined questions.
By Jiayao Chen, Shi Liu, Linyi Yang
arXiv:2607. 02329v1 Announce Type: new Abstract: Autonomous-research agents have demonstrated end-to-end LLM automation in machine-learning sandboxes where execution provides calibration.
By Haonan Huang