arXiv AI By Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel, Anjishnu Banerjee, Priya Deshpande, Bradley Taylor, Anai N. Kothari

Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

Read the original on arXiv AI →

The study examined biomedical research articles from 2022 to March 2026 that employed large language models (LLMs). It found that 42% of the most frequently used models were already retired or scheduled to retire within two years of publication, with a median retirement interval of 538 days. This high rate of model deprecation threatens the reproducibility of biomedical AI research.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 11

The widening evaluation gap in medical large language model research 2023 to 2026

The article examines how medical research lags behind the rapid evolution of large language models (LLMs). From January 2023 to June 2026, PubMed records in fourteen clinical domains grew 45‑fold, yet only 2.5 % employed randomized, controlled, or prospective designs. The evaluation gap widened from 1.33 to 6.08 quarters, with randomized trials assessing models that were on average 4.6 quarters older than other studies, and 62 % of such trials evaluated discontinued model families.

By Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif