arXiv AI

Harnessing agent memory to build lifelong AI partners for materials scientists

arXiv:2608. 11224v1 Announce Type: new Abstract: Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and judgement that links a new question to an old result.

arXiv AI
2d ago

CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

CompMat-Bench is a new benchmark comprising 94 tasks drawn from recent computational materials science studies, designed to evaluate AI agents on realistic research steps without requiring costly simulations during testing. The benchmark pre‑reproduces inputs and outputs to provide ground truth, allowing agents to be graded with fixed rules rather than an LLM judge. It supports both single tasks and multi‑step workflows, with varying levels of methodological guidance, and shows that while agents can achieve high pass rates on individual tasks, performance drops in longer workflows or with reduced guidance, often due to scientific rather than software errors.

By Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson
arXiv AI
Sep 3

Can Coding Agents Reproduce Findings in Computational Materials Science?

The paper introduces AutoMat, a benchmark designed to test large language model (LLM) coding agents on their ability to reproduce claims from computational materials science. AutoMat presents three challenges: reconstructing underspecified procedures, navigating specialized toolchains, and assessing whether the evidence supports a claim. Experiments show that current LLM agents achieve low success rates, with the best setting reaching only 53%, and failures stem mainly from incomplete procedures, methodological deviations, and execution fragility.

By Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, Wyatt Bunstine, William Jurayj, Somdatta Goswami, Tyrel McQueen, Michael Shields, Jaafar El-Awady, Paulette Clancy, Benjamin Van Durme, Nicholas Andrews, William Walden, Daniel Khashabi
Hugging Face Trending Papers
Sep 2

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills introduces DisCo, a research agent that extracts and verifies operational knowledge from GitHub repositories to create reusable AI skills. The agent produces both task‑agnostic skills—compiled into the AREX‑Skill Library of over 5,000 verified skills from 1,000 repositories—and task‑oriented skills tailored to specific research tasks. When equipped with these skills, the agent achieves significant performance gains across multiple benchmarks, outperforming a skill‑free version by 134.3% on MLE‑bench, 34.4% on PaperBench, 9.2% on FrontierCS, and 14.0% on PassNet.

arXiv AI
Sep 3

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

The paper introduces Repo-To-Skill, a method for converting GitHub repositories into reusable AI skills. By distilling operational knowledge from over 1,000 machine‑learning repositories, the authors build the AREX‑Skill Library with more than 5,000 verified skills across 20 areas. Integrating these skills into a research agent—DisCo—yields significant performance boosts on multiple benchmarks, demonstrating the value of reusable, task‑agnostic knowledge.

By Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen, Defu Lian, Zhicheng Dou, Chaozhuo Li, Qiwei Ye, Zheng Liu
arXiv AI
Sep 15

OpenAl4S: Code as Action, Science as Sessions

OpenAI4S is an open‑source scientific research agent that treats code as action and science as sessions, combining a persistent computing runtime with structured session management. It uses tool calls for orchestration, executes code cells in persistent Python and R kernels, and records an append‑only Action Ledger, per‑cell execution logs, versioned artifacts, environment snapshots, and workspace checkpoints to preserve provenance and enable session recovery, branching, and extension. Evaluated on 36 research scenarios—including retrosynthesis, molecular dynamics, and protein design—OpenAI4S achieved a higher overall score (7.83) than a general‑purpose coding harness, especially on long‑horizon, computation‑intensive workflows, though reproducibility remains an open challenge. whyItMatters":"The system demonstrates that persistent execution coupled with session‑level provenance can enhance the reliability of AI‑assisted scientific workflows, as evidenced by its superior performance across diverse research scenarios."

By Gongbo Zhang, Hao Li, Yu Wang, Mujie Lin, Liuzhenghao Lv, Yicheng Mao, Yimi Wang, Jun Zhu, Minhan Tang, Zhengxiang Jiang, Yusong Wang, Jiayu Yao, Kunpeng Ning, Dawei Pang, Yonghong Tian, OpenAI4S Community, Yuyang Liu, Li Yuan
arXiv AI
Jun 6

AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations

arXiv:2605. 26179v2 Announce Type: replace-cross Abstract: Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting algorithms when convergence stalls, revising plans when unexpected physics emerges, and inserting steps as intermediate results reshape the problem.

By Penghui Yang, Zhonghan Zhang, Yue Li, Xinrun Wang, Yanchen Deng, Yuhao Lu, Bijun Tang, Zheng Liu, Bo An
arXiv AI
Jun 12

Fantastic Scientific Agents and How to Build Them: AgentBuild for Rietveld Refinement

arXiv:2606. 12834v1 Announce Type: new Abstract: As scientific workflows shift from deterministic executables to LLM-based agents, the development practices on offer, such as fine-tuning, reinforcement learning, and prompt-and-go, bury the scientist's judgment.

By Woong Shin, Craig A. Bridges, Marshall T. McDonnell, Rafael Ferreira da Silva
arXiv AI
Aug 14

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

arXiv:2512. 19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation.

By Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Rui Ye, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen
arXiv AI
Jun 30

Hierarchical Experimentalist Agents

arXiv:2606. 29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search.

By Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka, Varun Gandhi, Scott Niekum