arXiv:2607. 24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally.
By Carlo Iacono (Charles Sturt University, Australia)
Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes.
The article examines how medical research lags behind the rapid evolution of large language models (LLMs). From January 2023 to June 2026, PubMed records in fourteen clinical domains grew 45‑fold, yet only 2.5 % employed randomized, controlled, or prospective designs. The evaluation gap widened from 1.33 to 6.08 quarters, with randomized trials assessing models that were on average 4.6 quarters older than other studies, and 62 % of such trials evaluated discontinued model families.
By Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
arXiv:2604. 12243v2 Announce Type: replace-cross Abstract: Identifying promising research directions in fast-moving subareas is one of the most cognitively expensive tasks in modern AI research.
By Jinkai Tao, Yubo Wang, Xiaoyu Liu, Menglin Yang
arXiv:2609.14988v1 Announce Type: new
Abstract: Background. Large language models are increasingly used to help write biomedical text but may fabricate references to nonexistent work. How often large...
By Maxim Topaz, Zhihong Zhang, Nir Roguin, Pallavi Gupta, Zichao Li, Laura-Maria Peltonen
arXiv:2606. 02184v1 Announce Type: cross Abstract: These names do not exist.
By Micha{\l} Brzozowski, Neo Christopher Chung