The study examined biomedical research articles from 2022 to March 2026 that employed large language models (LLMs). It found that 42% of the most frequently used models were already retired or scheduled to retire within two years of publication, with a median retirement interval of 538 days. This high rate of model deprecation threatens the reproducibility of biomedical AI research.
By Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel, Anjishnu Banerjee, Priya Deshpande, Bradley Taylor, Anai N. Kothari
arXiv:2607. 24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally.
By Carlo Iacono (Charles Sturt University, Australia)
Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes.
The article examines how medical research lags behind the rapid evolution of large language models (LLMs). From January 2023 to June 2026, PubMed records in fourteen clinical domains grew 45‑fold, yet only 2.5 % employed randomized, controlled, or prospective designs. The evaluation gap widened from 1.33 to 6.08 quarters, with randomized trials assessing models that were on average 4.6 quarters older than other studies, and 62 % of such trials evaluated discontinued model families.
By Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
The paper introduces SciUtopia, a closed‑loop large‑language‑model simulation framework that models the evolving academic research ecosystem, including research direction, collaboration, publication, peer review, funding, and researcher attrition. Running 61 simulation worlds, the system simulates over 40,000 researchers and 1.2 million LLM‑generated peer reviews, revealing that rejection‑driven resubmission increases reviewer burden, cautious exploration balances citation impact with career success, and resource inequality can arise without early‑funding advantage.
By Yiqiao Jin, Yiyang Wang, Lucheng Fu, Bing He, Siheng Xiong, Yijia Xiao, B. Aditya Prakash, Josiah Hester, Srijan Kumar, James Evans, Jindong Wang
arXiv:2607. 21327v1 Announce Type: cross Abstract: Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, technology, and innovation (STI) analytics, yet suffer from temporal lag, semantic shallowness, and an inability to capture the non-linear dynamics of contemporary knowledge ecosystems.
By Muhsen Hammoud
arXiv:2606. 26130v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear.
By Francesca Carlon, Brecht Verbeken, Vincent Ginis, Andres Algaba
arXiv:2608. 09964v1 Announce Type: cross Abstract: The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer Society since 2012 and publishing work from institutions across the country.
By Thales Sales Almeida, Giovana Kerche Bon\'as, Thiago Laitz, Jo\~ao Guilherme Alves Santos, Hugo Abonizio, Roseval Malaquias Junior, Marcos Piau, Celio Larcher, Ramon Pires, Rodrigo Nogueira
arXiv:2605. 04135v2 Announce Type: replace-cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do.
By David Gringras, Misha Salahshoor
arXiv:2606. 24901v1 Announce Type: new Abstract: Continual learning capability is critical for Industrial LLMs, as deployed models must be continuously updated to meet evolving requirements and environments, rather than repeatedly retrained from scratch.
By Hao Jiang, Enneng Yang, Guojie Zhu, Yibin Chen, Yunkun Xu, Zifu Kou, Jiayi Li, Chong Chen, Zhao Cao, Li Shen
The study analyzes 207,111 astronomy papers from 2015 to mid‑2026 to quantify how many contain language‑model‑generated vocabulary. Using a hierarchical Bayesian model calibrated on pre‑2020 unassisted papers and 392 papers that disclose model use, the authors estimate that in 2025 roughly 54% (±8% statistical, ±26% systematic) of papers show a language‑model trace, with the estimate remaining above 36% under various assumptions. Despite only 0.81% of 2025 papers explicitly declaring model assistance, the trace is pervasive, and the detectable signal is fading as authors adapt to the characteristic words.
whyItMatters":"The findings reveal that language‑model assistance has become widespread in recent astronomy research, yet most authors do not disclose its use, highlighting a growing gap between actual practice and transparency in scholarly writing."
By Serat M. Saad, Yuan-Sen Ting
arXiv:2607. 28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect.
By Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin