arXiv:2607. 09762v1 Announce Type: new Abstract: Public battery aging datasets are a critical asset for advanced health management, but their practical use is often limited by inconsistent formats, unclear schemas, and metadata scattered across repositories and publications.
By Tianwen Zhu, Hao Wang, Yonggang Wen
OpenAI4S is an open‑source scientific research agent that treats code as action and science as sessions, combining a persistent computing runtime with structured session management. It uses tool calls for orchestration, executes code cells in persistent Python and R kernels, and records an append‑only Action Ledger, per‑cell execution logs, versioned artifacts, environment snapshots, and workspace checkpoints to preserve provenance and enable session recovery, branching, and extension. Evaluated on 36 research scenarios—including retrosynthesis, molecular dynamics, and protein design—OpenAI4S achieved a higher overall score (7.83) than a general‑purpose coding harness, especially on long‑horizon, computation‑intensive workflows, though reproducibility remains an open challenge.
whyItMatters":"The system demonstrates that persistent execution coupled with session‑level provenance can enhance the reliability of AI‑assisted scientific workflows, as evidenced by its superior performance across diverse research scenarios."
By Gongbo Zhang, Hao Li, Yu Wang, Mujie Lin, Liuzhenghao Lv, Yicheng Mao, Yimi Wang, Jun Zhu, Minhan Tang, Zhengxiang Jiang, Yusong Wang, Jiayu Yao, Kunpeng Ning, Dawei Pang, Yonghong Tian, OpenAI4S Community, Yuyang Liu, Li Yuan
arXiv:2607. 16038v1 Announce Type: new Abstract: Scientific work increasingly spans heterogeneous artifacts -- papers, code, datasets, scientific file formats, model outputs, figures, manuscripts, and team decisions -- yet general-purpose AI assistants rarely preserve these objects as a coherent, auditable research state.
By SciForge Team, Zhangyang Gao, Minghao Fang, Yifei Liu, Hanhui Yang, Xinyu Gu, Shixiang Tang, Siqi Sun, Lei Bai, Cheng Tan, Mengdi Liu, Hao Wu, Shuizhou Chen
The paper extends Co‑Scientist, a Gemini‑based multi‑agent system, and validates it in real‑world scientific settings. In materials science it designed a safe precursor route for MXenes and achieved single‑attempt growth of monolayer MoS₂, MoSe₂, and WS₂. In biology it predicted swarming phenotypes of engineered E. coli, and in computer science it discovered a superior inference‑time scaling architecture for HealthBench. A double‑blind study with 30 experts showed that Co‑Scientist’s reliability modules reduce hallucination and plagiarism while improving research safety.
By Samuel Schmidgall, Xiaokai Zhu, Marian Shaw, Lin Yang, Valentin Li\'{e}vin, Jingyun Yang, Yuchen Zhuang, Tim Strother, Alex Bijamov, Min Woo Sun, Anil Palepu, Justin Chen, David Steiner, Jacqueline Shreibati, Wei-Hung Weng, Yilin Zhao, Xingjian Hu, Nicholas Zahn, Sadhya Garg, Julia Kirby, Yuxiang Gan, Jiaoli Li, Divy Thakkar, Shekoofeh Azizi, David Racz, Juraj Gottweis, Vivek Natarajan, Chenglin Wu, Tal Danino, Keran Rong, Haozhe Wang, Benoit Schillings, Yong Cheng, Quoc V. Le, Tao Tu
AutoREC is an open‑source Python platform that uses reinforcement learning to automatically generate equivalent circuit models (ECMs) from electrochemical impedance spectroscopy (EIS) data. The platform frames ECM generation as a Markov decision process, training a Double Deep Q‑Network agent that iteratively modifies circuit topologies based on state, actions, and model feedback. It supports end‑to‑end workflows—including EIS preprocessing, agent training, ECM generation, and visualization—and has been demonstrated on synthetic datasets and real experimental spectra from batteries, corrosion, oxygen evolution, and CO₂ reduction systems.
By Ali Jaberi (Clean Energy Innovation Research Centre, National Research Council Canada, Mississauga, ON, Canada), Yonatan Kurniawan (Department of Materials Science and Engineering, University of Toronto, Toronto, ON, Canada), Robert Black (Clean Energy Innovation Research Centre, National Research Council Canada, Mississauga, ON, Canada), Shayan Mousavi M. (Clean Energy Innovation Research Centre, National Research Council Canada, Mississauga, ON, Canada), Kabir Verma (Cheriton School of Computer Science, University of Waterloo, Waterloo, ON, Canada), Zoya Sadighi (Clean Energy Innovation Research Centre, National Research Council Canada, Mississauga, ON, Canada), Santiago Miret (Lila Sciences, San Francisco, CA, USA), Jason Hattrick-Simpers (Department of Materials Science and Engineering, University of Toronto, Toronto, ON, Canada)
arXiv:2607. 23886v1 Announce Type: cross Abstract: X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature.
By Tanjin He, Aikaterini Vriza, Logan Ward, Xu Huang, Yiming Chen, Anubhav Jain, Gerbrand Ceder, Rajeev S. Assary, Ian T. Foster, Maria K. Y. Chan
The article reviews how Large Models (LMs) based on Transformer architectures and self‑supervised pre‑training can address longstanding challenges in Battery Prognostics and Health Management (BPHM). It surveys LM applications across data scarcity, generalization, interpretability, and system automation, and outlines a roadmap for future research, including collaborative data ecosystems, validation, trustworthiness, and efficient deployment. The review aims to guide researchers and practitioners in developing next‑generation battery management systems that are safe, reliable, and autonomous throughout battery lifecycles.
By Jiale Liu, Huan Wang, Weicheng Wang, Rong Zhu, Qiqi Wang, Min Xie
ProtoMI is a literature‑driven framework that learns structural priors from 126 reported boron‑containing electrolyte additives and applies them to screen 179,977 unlabeled candidates. Using graph contrastive learning, it identifies seven interpretable prototypes and adapts them through semi‑supervised contrastive learning, achieving enrichment factors of 9.2–45.6 while evaluating less than 2% of the candidate space. The method led to the discovery of four commercially accessible additives, including TNDB, which improves high‑temperature LiFePO4||graphite cycling by 34.93% and forms protective interphases that suppress solvent decomposition and Fe deposition.
By Weixiang Hong, Hongting Du, Jiayue Tang, Ruifeng Tan, Yangjian Quan, Jia Li, Jiaqiang Huang
arXiv:2608. 04942v1 Announce Type: cross Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications.
By Brendan Smith, Susana Lopez-Moreno, Eric Dolores-Cuenca, Sangil Kim, Jose L. Mendoza-Cortes, Nijamudheen Abdulrahiman
arXiv:2603. 01421v3 Announce Type: replace Abstract: While large language models accelerate scientific discovery, existing agents face severe limitations in adaptability, domain generalization, and multimodal scalability, often struggling to autonomously process raw, domain-specific experimental data.
By Ke Lin, Owais Aijaz, Yilin Lu, Yiyang Luo, Xuehang Guo, Preslav Nakov
arXiv:2607. 11079v1 Announce Type: new Abstract: Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria.
By Chuhan Shi, Xiaoquan Ren, Sicheng Song, Haobo Li, Rui Sheng, Yushi Sun
arXiv:2608.31076v1 Announce Type: cross
Abstract: Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experi...
By Xuehai Wang, Haowei Qin, Tongxin Liu, Junkai Li, Buqiang Xu, Jintian Zhang, Yijun Chen, Zirui Xue, Shumin Deng