arXiv:2605.11280v2 Announce Type: replace-cross
Abstract: Fast surrogate models for expensive simulations are now essential across the sciences, yet they typically operate as black boxes. We present...
By Tousif Islam, Digvijay Wadekar, Tejaswi Venumadhav, Matias Zaldarriaga, Ajit Kumar Mehta, Javier Roulet, Barak Zackay
arXiv:2605.11269v2 Announce Type: replace-cross
Abstract: Modern gravitational wave astronomy relies on modeling tasks that often require months of graduate-level effort, including building fast wave...
By Tousif Islam, Digvijay Wadekar, Zihan Zhou
The paper introduces a framework for evaluating large language model agents by attempting to end‑to‑end reproduce published astronomy studies, separating execution from verification and distinguishing computational failures from methodological ambiguities. Applying this to fourteen papers—one from The Astrophysical Journal and thirteen from Nature—revealed that eleven contained ambiguities that prevented a uniquely specified reproduction path. In a controlled case study, twelve different analysis paths produced distance estimates ranging from 2.16 to 3.53 kpc, with only one matching the published value of ~2.70 kpc, demonstrating that matching outcomes does not guarantee that the agent has reconstructed the underlying reasoning.
whyItMatters":"The study shows that end‑to‑end reproduction can expose gaps in implicit scientific knowledge within AI systems, highlighting the need for better integration of causal relevance in LLM agents."
By Yuehui Wang, Xinyu Qi, Guirong Xue, Cheng Wang, Yangbin Xie, Xiaoyu Tang, Cong Sun
arXiv:2609.27490v1 Announce Type: new
Abstract: AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding...
By Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component cha...
arXiv:2607. 01372v1 Announce Type: cross Abstract: Gravitational Waves (GWs) represent the newest window of astronomy, furthering our understanding of compact objects like black holes and neutron stars in the Universe.
By Bhavya Gupta, Deep Chatterjee, William Benoit, Ethan Marx, Christina Reissel, Seiya Tsukamoto, Kyungseop Yoon, Michael W. Coughlin, Philip Harris, Erik Katsavounidis
arXiv:2605. 14791v2 Announce Type: replace-cross Abstract: Recent advances in artificial intelligence (AI) agents are pushing AI beyond tools toward autonomous scientific discovery.
By Licong Xu, Thomas Borrett
arXiv:2607. 25145v1 Announce Type: cross Abstract: We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond.
By Takuya Isogawa, Ryotaro Okabe, Nutdech Phadetsuwannukun, Mingda Li, Paola Cappellaro
La Agente ’Optima is an agentic framework that builds and manages Bayesian optimization campaigns for self‑driving laboratories, separating large language model reasoning from campaign execution. It maintains a persistent optimization state, allowing consistent repetitive loops and auditable decisions, and only returns control to the agent when interpretation or revision is needed. In tests on digital discovery tasks and physical platforms, it corrected measurement failures, improved yields, and recommended formulation changes, outperforming human‑directed campaigns in cost and material usage.
By Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garc\'ia Carrillo, Yeonghun Kang, Juan B. P\'erez-S\'anchez, Simone Pilon, Martin Fitzner, Timothy No\"el, Frank Gu, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv:2609.12533v1 Announce Type: cross
Abstract: Real-world Earth observation (EO) agents must translate high-level scientific questions into executable workflows to acquire observations, prepare da...
By Zhutao Lv, Chenhao Dang, Yi Feng, Yanpei Gong, Xiaolei Wang, Junyan Ye, Conghui He, Weijia Li
arXiv:2609.24165v1 Announce Type: new
Abstract: Synchrotron data reduction, detector calibration followed by azimuthal integration of terabyte-scale diffraction series, is a multi-step, expert-bound...
By Pawan K. Tripathi, Hemant Sharma, Andrew Chuang, Mathew J. Cherukara
arXiv:2608.28590v1 Announce Type: new
Abstract: Large Language Model (LLM) agents have shown promise for automating data-science workflows, yet their end-to-end performance depends critically on the...
By Fan Liu, Hao Liu