arXiv AI By Yuehui Wang, Xinyu Qi, Guirong Xue, Cheng Wang, Yangbin Xie, Xiaoyu Tang, Cong Sun

Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy

Read the original on arXiv AI →

The paper introduces a framework for evaluating large language model agents by attempting to end‑to‑end reproduce published astronomy studies, separating execution from verification and distinguishing computational failures from methodological ambiguities. Applying this to fourteen papers—one from The Astrophysical Journal and thirteen from Nature—revealed that eleven contained ambiguities that prevented a uniquely specified reproduction path. In a controlled case study, twelve different analysis paths produced distance estimates ranging from 2.16 to 3.53 kpc, with only one matching the published value of ~2.70 kpc, demonstrating that matching outcomes does not guarantee that the agent has reconstructed the underlying reasoning. whyItMatters":"The study shows that end‑to‑end reproduction can expose gaps in implicit scientific knowledge within AI systems, highlighting the need for better integration of causal relevance in LLM agents."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 8

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

arXiv:2607. 05682v1 Announce Type: new Abstract: LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect.

By Yufeng Wang
arXiv AI
Sep 7

La Agente \'Optima: Towards Agentic Self-Driving Laboratories

La Agente ’Optima is an agentic framework that builds and manages Bayesian optimization campaigns for self‑driving laboratories, separating large language model reasoning from campaign execution. It maintains a persistent optimization state, allowing consistent repetitive loops and auditable decisions, and only returns control to the agent when interpretation or revision is needed. In tests on digital discovery tasks and physical platforms, it corrected measurement failures, improved yields, and recommended formulation changes, outperforming human‑directed campaigns in cost and material usage.

By Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garc\'ia Carrillo, Yeonghun Kang, Juan B. P\'erez-S\'anchez, Simone Pilon, Martin Fitzner, Timothy No\"el, Frank Gu, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv AI
Aug 25

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

arXiv:2608.23525v1 Announce Type: new Abstract: Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards ma...

By Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong, Weidong Tang, Yuxuan Liang, Weijia Li, Ming Jin, Shirui Pan, Yuhao Kang, Dingyi Zhuang, Jinhua Zhao