arXiv AI

Agentic AI for Gravitational Wave Data Analysis: A Head-to-Head Comparison of Coding Agents Executing a Matched Filter Pipeline on Einstein Telescope Simulated Data

arXiv AI
4d ago

Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy

The paper introduces a framework for evaluating large language model agents by attempting to end‑to‑end reproduce published astronomy studies, separating execution from verification and distinguishing computational failures from methodological ambiguities. Applying this to fourteen papers—one from The Astrophysical Journal and thirteen from Nature—revealed that eleven contained ambiguities that prevented a uniquely specified reproduction path. In a controlled case study, twelve different analysis paths produced distance estimates ranging from 2.16 to 3.53 kpc, with only one matching the published value of ~2.70 kpc, demonstrating that matching outcomes does not guarantee that the agent has reconstructed the underlying reasoning. whyItMatters":"The study shows that end‑to‑end reproduction can expose gaps in implicit scientific knowledge within AI systems, highlighting the need for better integration of causal relevance in LLM agents."

By Yuehui Wang, Xinyu Qi, Guirong Xue, Cheng Wang, Yangbin Xie, Xiaoyu Tang, Cong Sun
arXiv AI
Jul 3

AI-enabled gravitational-waves searches for binary neutron stars at optimal sensitivity

arXiv:2607. 01372v1 Announce Type: cross Abstract: Gravitational Waves (GWs) represent the newest window of astronomy, furthering our understanding of compact objects like black holes and neutron stars in the Universe.

By Bhavya Gupta, Deep Chatterjee, William Benoit, Ethan Marx, Christina Reissel, Seiya Tsukamoto, Kyungseop Yoon, Michael W. Coughlin, Philip Harris, Erik Katsavounidis
arXiv AI
Sep 7

La Agente \'Optima: Towards Agentic Self-Driving Laboratories

La Agente ’Optima is an agentic framework that builds and manages Bayesian optimization campaigns for self‑driving laboratories, separating large language model reasoning from campaign execution. It maintains a persistent optimization state, allowing consistent repetitive loops and auditable decisions, and only returns control to the agent when interpretation or revision is needed. In tests on digital discovery tasks and physical platforms, it corrected measurement failures, improved yields, and recommended formulation changes, outperforming human‑directed campaigns in cost and material usage.

By Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garc\'ia Carrillo, Yeonghun Kang, Juan B. P\'erez-S\'anchez, Simone Pilon, Martin Fitzner, Timothy No\"el, Frank Gu, Varinia Bernales, Al\'an Aspuru-Guzik