The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing
arXiv:2606. 02184v1 Announce Type: cross Abstract: These names do not exist.
These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived.
arXiv:2606. 02184v1 Announce Type: cross Abstract: These names do not exist.
arXiv:2609.15369v1 Announce Type: new Abstract: Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-l...
arXiv:2608. 05157v1 Announce Type: cross Abstract: Double blind peer review serves as the scientific community primary defense against status and affiliation bias.
The study examined biomedical research articles from 2022 to March 2026 that employed large language models (LLMs). It found that 42% of the most frequently used models were already retired or scheduled to retire within two years of publication, with a median retirement interval of 538 days. This high rate of model deprecation threatens the reproducibility of biomedical AI research.
arXiv:2606. 11105v1 Announce Type: cross Abstract: Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them.
The paper audits a 366‑day autobiographical book generated by a large language model (LLM) against an independent verification corpus. Using a four‑level rubric, 354 of the 366 days (96.7%) failed verification, with only 12 days containing corroborated scenes and 19 days containing actively contradicted claims. Regenerating the same days with current models yielded 100% verification failure, while grounding the generation in the subject’s own corpus improved the rate to 83.3% but still left substantial residual failure.
arXiv:2609.14988v1 Announce Type: new Abstract: Background. Large language models are increasingly used to help write biomedical text but may fabricate references to nonexistent work. How often large...
INDRA is a research platform that integrates multiple archival collections—such as UCSF’s Industry Documents Library, Columbia and CUNY’s ToxicDocs, and Stanford’s SRITA—into a single, LLM‑readable corpus. It employs three safeguards: a closed evidentiary sandbox, real‑time provenance tagging, and a deterministic system‑level protocol to ensure that model outputs are clearly distinguished from archival evidence and from the model’s own inferences. The platform enables large‑language‑model‑powered investigations across these archives while keeping the conditions of knowledge production transparent and auditable.
arXiv:2606. 10794v1 Announce Type: new Abstract: As agentic applications increasingly route user tasks through official and third-party LLM APIs, provenance becomes an operational question: which model generated a given black-box response?
arXiv:2606. 18060v1 Announce Type: new Abstract: As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important.
arXiv:2607. 00738v1 Announce Type: cross Abstract: Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival record.
Large language models are increasingly deployed in citation-augmented settings, yet the effect of citation presence on model behavior independent of factual content remains poorly understood. We introduce AuthorityBench, a 220,564-prompt multi-domain benchmark that isolates how citation-based authority signals influence epistemic behavior in LLMs.