arXiv:2608. 10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty.
By Sanidhya Vijayvargiya, Rahul Lokesh
arXiv:2609.38962v1 Announce Type: new
Abstract: Recent work on hallucination detection in large language models has shown that, for a fixed pre-trained model and reasoning task, it is possible to est...
By Litian Liu, Qiqi Hou, Yubing Jian, Reza Pourreza, Mohammad Ghavamzadeh, Roland Memisevic, Yao Qin, Hong Cai
arXiv:2507. 06722v2 Announce Type: replace-cross Abstract: Understanding how large language models (LLMs) internally represent and process their predictions is central to detecting uncertainty and preventing hallucinations.
By Sunwoo Kim, Haneul Yoo, Alice Oh
arXiv:2607. 17883v1 Announce Type: cross Abstract: Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true.
By Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu
arXiv:2609.35804v1 Announce Type: cross
Abstract: Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment...
By Mamehgol Yousefi, Ahmad Shahi, Mos Sharifi, Alvaro Romera, Simon Hoermann, Tham Piumsomboon
arXiv:2601.22984v3 Announce Type: replace
Abstract: Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evalu...
By Yuhao Zhan, Tianyu Fan, Linxuan Huang, Zirui Guo, Chao Huang