arXiv AI By Carlo Iacono (Charles Sturt University, Australia)

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

Read the original on arXiv AI →

arXiv:2607. 24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

The study examined biomedical research articles from 2022 to March 2026 that employed large language models (LLMs). It found that 42% of the most frequently used models were already retired or scheduled to retire within two years of publication, with a median retirement interval of 538 days. This high rate of model deprecation threatens the reproducibility of biomedical AI research.

By Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel, Anjishnu Banerjee, Priya Deshpande, Bradley Taylor, Anai N. Kothari
arXiv AI
Aug 28

6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation

The paper introduces a six‑stage audit framework for assessing reproducibility in computer science literature and applies it to the neuro‑symbolic AI (NSAI) subfield. Using the framework, the authors screened 5,497 records, identified 1,304 eligible studies, and found verifiable code artifacts for only 455 of them. Of those, they fully or partially reproduced 85 studies, representing 6.52% of the eligible corpus and 18.68% of attempted reruns, highlighting a significant reproducibility gap even when code is declared available.

By Brandon Colelough, Vladimir Martirosyan, Ishan Tamrakar, William Regli, Aditya Kumar, Anh N. Nhu, Dhruv Dubey, Raj Ambavane, Haowei Deng