arXiv:2606. 11105v1 Announce Type: cross Abstract: Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them.
By Haeji Jung, Hila Gonen
arXiv:2601.22984v3 Announce Type: replace
Abstract: Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evalu...
By Yuhao Zhan, Tianyu Fan, Linxuan Huang, Zirui Guo, Chao Huang
In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration...
arXiv:2603. 09986v3 Announce Type: replace-cross Abstract: Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious problem within natural language processing for which we do not yet have an effective solution to mitigate against.
By Brandon C. Colelough, Davis Bartels, Dina Demner-Fushman
arXiv:2606. 03628v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in open-ended text generation, yet they remain prone to hallucinating incorrect or unsupported content, which undermines their reliability.
By Lin Li, Georgia Channing, Suhaas M Bhat, Gabriel Davis Jones, Yarin Gal
arXiv:2606. 07537v1 Announce Type: cross Abstract: Large language models hallucinate--producing fluent, confident, factually wrong outputs--with a consistency that persists across generations and scales.
By Md. Rejaul Korim Sadi, Toufiqur Rahman Tasin, Golam Mostofa Naeem