arXiv:2607. 24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally.
By Carlo Iacono (Charles Sturt University, Australia)
arXiv:2605. 04135v2 Announce Type: replace-cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do.
By David Gringras, Misha Salahshoor
The study examined biomedical research articles from 2022 to March 2026 that employed large language models (LLMs). It found that 42% of the most frequently used models were already retired or scheduled to retire within two years of publication, with a median retirement interval of 538 days. This high rate of model deprecation threatens the reproducibility of biomedical AI research.
By Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel, Anjishnu Banerjee, Priya Deshpande, Bradley Taylor, Anai N. Kothari
arXiv:2608. 05179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review.
By Tianyu Ding, Aditya Nannapaneni, Bingfan Liu, Ling Zhang
The paper introduces a six‑stage audit framework for assessing reproducibility in computer science literature and applies it to the neuro‑symbolic AI (NSAI) subfield. Using the framework, the authors screened 5,497 records, identified 1,304 eligible studies, and found verifiable code artifacts for only 455 of them. Of those, they fully or partially reproduced 85 studies, representing 6.52% of the eligible corpus and 18.68% of attempted reruns, highlighting a significant reproducibility gap even when code is declared available.
By Brandon Colelough, Vladimir Martirosyan, Ishan Tamrakar, William Regli, Aditya Kumar, Anh N. Nhu, Dhruv Dubey, Raj Ambavane, Haowei Deng
arXiv:2604. 12243v2 Announce Type: replace-cross Abstract: Identifying promising research directions in fast-moving subareas is one of the most cognitively expensive tasks in modern AI research.
By Jinkai Tao, Yubo Wang, Xiaoyu Liu, Menglin Yang