arXiv AI

The Shrinking Lifespan of LLMs in Science

arXiv:2604. 07530v2 Announce Type: replace-cross Abstract: Scaling laws describe how language model capabilities grow with compute and data, but say nothing about how long a model matters once released.

arXiv AI
Sep 7

Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

The study examined biomedical research articles from 2022 to March 2026 that employed large language models (LLMs). It found that 42% of the most frequently used models were already retired or scheduled to retire within two years of publication, with a median retirement interval of 538 days. This high rate of model deprecation threatens the reproducibility of biomedical AI research.

By Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel, Anjishnu Banerjee, Priya Deshpande, Bradley Taylor, Anai N. Kothari
arXiv Computation and Language
Sep 11

The widening evaluation gap in medical large language model research 2023 to 2026

The article examines how medical research lags behind the rapid evolution of large language models (LLMs). From January 2023 to June 2026, PubMed records in fourteen clinical domains grew 45‑fold, yet only 2.5 % employed randomized, controlled, or prospective designs. The evaluation gap widened from 1.33 to 6.08 quarters, with randomized trials assessing models that were on average 4.6 quarters older than other studies, and 62 % of such trials evaluated discontinued model families.

By Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
arXiv AI
2d ago

Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems

The paper introduces SciUtopia, a closed‑loop large‑language‑model simulation framework that models the evolving academic research ecosystem, including research direction, collaboration, publication, peer review, funding, and researcher attrition. Running 61 simulation worlds, the system simulates over 40,000 researchers and 1.2 million LLM‑generated peer reviews, revealing that rejection‑driven resubmission increases reviewer burden, cautious exploration balances citation impact with career success, and resource inequality can arise without early‑funding advantage.

By Yiqiao Jin, Yiyang Wang, Lucheng Fu, Bing He, Siheng Xiong, Yijia Xiao, B. Aditya Prakash, Josiah Hester, Srijan Kumar, James Evans, Jindong Wang
arXiv AI
Jul 24

From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics

arXiv:2607. 21327v1 Announce Type: cross Abstract: Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, technology, and innovation (STI) analytics, yet suffer from temporal lag, semantic shallowness, and an inability to capture the non-linear dynamics of contemporary knowledge ecosystems.

By Muhsen Hammoud
arXiv AI
Aug 12

Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems

arXiv:2608. 09964v1 Announce Type: cross Abstract: The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer Society since 2012 and publishing work from institutions across the country.

By Thales Sales Almeida, Giovana Kerche Bon\'as, Thiago Laitz, Jo\~ao Guilherme Alves Santos, Hugo Abonizio, Roseval Malaquias Junior, Marcos Piau, Celio Larcher, Ramon Pires, Rodrigo Nogueira
arXiv Machine Learning
Jun 25

LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning

arXiv:2606. 24901v1 Announce Type: new Abstract: Continual learning capability is critical for Industrial LLMs, as deployed models must be continuously updated to meet evolving requirements and environments, rather than repeatedly retrained from scratch.

By Hao Jiang, Enneng Yang, Guojie Zhu, Yibin Chen, Yunkun Xu, Zifu Kou, Jiayi Li, Chong Chen, Zhao Cao, Li Shen
arXiv Computation and Language
Sep 11

More than half of recent astronomy papers are written with language-model assistance

The study analyzes 207,111 astronomy papers from 2015 to mid‑2026 to quantify how many contain language‑model‑generated vocabulary. Using a hierarchical Bayesian model calibrated on pre‑2020 unassisted papers and 392 papers that disclose model use, the authors estimate that in 2025 roughly 54% (±8% statistical, ±26% systematic) of papers show a language‑model trace, with the estimate remaining above 36% under various assumptions. Despite only 0.81% of 2025 papers explicitly declaring model assistance, the trace is pervasive, and the detectable signal is fading as authors adapt to the characteristic words. whyItMatters":"The findings reveal that language‑model assistance has become widespread in recent astronomy research, yet most authors do not disclose its use, highlighting a growing gap between actual practice and transparency in scholarly writing."

By Serat M. Saad, Yuan-Sen Ting
arXiv AI
Aug 3

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

arXiv:2607. 28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect.

By Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin