arXiv:2608. 16353v1 Announce Type: cross Abstract: Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments.
By Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye, Junwei Zhang, Weiran Yao, Zhiwei Liu, Qingsong Wen, Yilei Shao
arXiv:2606. 01923v1 Announce Type: cross Abstract: Large Language Models (LLMs) frequently exhibit "contextual disregard" when faced with input evidence that conflicts with their internal parametric memory, leading to persistent factual hallucinations.
By Mingkuan Zhao, Yide Gao, Wentao Hu, Suquan Chen, Tianchen Huang, Zhenhua An, Zetao Chang, Xiayu Sun, Yuheng Min
arXiv:2608.16353v2 Announce Type: replace
Abstract: Internal-state probes enable truthfulness prediction before a large language model generates an answer. When detectors change both the layers they...
By Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye, Junwei Zhang, Weiran Yao, Zhiwei Liu, Qingsong Wen, Yilei Shao
arXiv:2609.00231v1 Announce Type: new
Abstract: Existing research on object hallucination in multimodal large language models (MLLMs) predominantly attributes the problem to language priors such as o...
By Peiyang Xu, Xiaopei Zhu, Jun Zhu, Xiaolin Hu
The paper proposes using the temporal volatility of internal attention mechanisms—measured by an unsupervised attention dispersion metric—as a diagnostic signal for hallucinations in large language models. It demonstrates that spikes in attention entropy within intermediate layers correlate with reasoning breakdowns, and shows statistically significant AUC improvements of up to +0.076 over output-based baselines on GSM8K and MATH-500 benchmarks using the Qwen2.5 model family.
By Shardul P. More, Tanuja S. Pawar
The paper investigates how the topology of attention graphs can differentiate hallucinated from non-hallucinated responses in large language models. By analyzing Forman-Ricci curvature, the authors identify structural bottlenecks and develop a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinations. Extensive evaluation across multiple LLMs and benchmarks shows that this single-pass approach consistently outperforms existing attention-based and multi-response baselines, while also revealing that impaired context sharing—such as over-reliance on self-attention and information over-squashing—correlates strongly with hallucination occurrences.
By Amir Jalilifard, Anderson Rocha, Eric Wong, Marcos Medeiros Raimundo
The paper introduces a multi‑signal pipeline for detecting hallucinations in large language models, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves high performance (F1 = 0.915, AUROC = 0.977) across QA, summarization, and dialogue, and shows that 25 % of training data yields 77 % of full‑data performance. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator cuts hallucination rates from 85.5 % to 37.7 %, and that domain‑specific fine‑tuning (PubMedBERT on SciFact) outperforms general‑domain models for biomedical text.
By Varun Teja Chundru, Debasmita Biswas
The paper introduces a multi‑signal pipeline for detecting hallucinations in large language model outputs, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves strong performance (F1 = 0.915, AUROC = 0.977) and further improves accuracy to 93.2% with MC Dropout. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator reduces hallucination rates from 85.5% to 37.7%, and show that domain‑specific fine‑tuning (PubMedBERT on SciFact) yields better results than general‑domain training.
arXiv:2606. 29431v1 Announce Type: new Abstract: Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image.
By Yichen Guo, Kai Tang, Fenglai Lin, Yiding Sun, Dongshuo Zhang, Wenya Wang, Lin William Cong, Shanghang Zhang
arXiv:2606. 07537v1 Announce Type: cross Abstract: Large language models hallucinate--producing fluent, confident, factually wrong outputs--with a consistency that persists across generations and scales.
By Md. Rejaul Korim Sadi, Toufiqur Rahman Tasin, Golam Mostofa Naeem
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
The paper investigates why vision‑language models that tokenize images with vector‑quantized (VQ) codebooks frequently hallucinate objects on grounded yes/no tasks. By applying activation patching across 25 models from eight large‑language‑model families, the authors uncover an early‑layer attention routing circuit shared by VQ‑tokenized VLMs. They develop a three‑gate diagnostic that isolates ten models carrying this circuit, show that swapping a single architectural component (VQ+Linear) introduces the circuit, and demonstrate that ablating the early‑layer ($L_0$) component reduces hallucinations in open‑ended generation by 31 % while other decoding‑time fixes do not.
By Shamanthak Hegde, Xiangrui Liu, Maitreya Patel, Yezhou Yang