Biomedical Reference Generation Remains Unreliable across 26 Large Language Models
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2607. 18360v1 Announce Type: cross Abstract: Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's accepted set.
The study examined biomedical research articles from 2022 to March 2026 that employed large language models (LLMs). It found that 42% of the most frequently used models were already retired or scheduled to retire within two years of publication, with a median retirement interval of 538 days. This high rate of model deprecation threatens the reproducibility of biomedical AI research.
arXiv:2606. 09500v1 Announce Type: new Abstract: Objective.
arXiv:2608. 11415v1 Announce Type: cross Abstract: Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists.
arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.
arXiv:2608. 19230v1 Announce Type: cross Abstract: As language models move from drafting prose to running literature-search agents with tool calls, fabricated references are becoming easier to catch and constrain.