arXiv Computation and Language

Biomedical Reference Generation Remains Unreliable across 26 Large Language Models

arXiv AI
Sep 7

Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

The study examined biomedical research articles from 2022 to March 2026 that employed large language models (LLMs). It found that 42% of the most frequently used models were already retired or scheduled to retire within two years of publication, with a median retirement interval of 538 days. This high rate of model deprecation threatens the reproducibility of biomedical AI research.

By Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel, Anjishnu Banerjee, Priya Deshpande, Bradley Taylor, Anai N. Kothari
arXiv AI
Jun 2

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.

By Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu
arXiv AI
Aug 28

MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection

MedFabric is a new benchmark for detecting word‑level medical fabrications, comprising 646 fabricated statements each paired with a ground‑truth passage that shares the same LLM authorship and nearly identical wording. The study shows that current detectors perform poorly—expert clinicians achieve only 53.3% macro‑F1 and no detector family surpasses 60% without gold evidence—highlighting that detection hinges on evidence correctness rather than subtlety of fabrication. The authors demonstrate that a retrieval‑confidence gate can substantially improve performance, raising macro‑F1 from 61% to 74%.

By Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin, Dongxu Zhang, Jun Han, Zhichao Yang, Davin Hill, Tamer Soliman, Sanjit Singh Batra, Robert Tillman, Guang Cheng