arXiv Computation and Language

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

arXiv AI
Jun 16

Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification

arXiv:2603. 24058v2 Announce Type: replace-cross Abstract: Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in high-stakes scenarios such as autonomous driving and medical image analysis.

By Han Sun, Qin Li, Peixin Wang, Min Zhang
arXiv Machine Learning
Sep 17

Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models

The paper proposes using the temporal volatility of internal attention mechanisms—measured by an unsupervised attention dispersion metric—as a diagnostic signal for hallucinations in large language models. It demonstrates that spikes in attention entropy within intermediate layers correlate with reasoning breakdowns, and shows statistically significant AUC improvements of up to +0.076 over output-based baselines on GSM8K and MATH-500 benchmarks using the Qwen2.5 model family.

By Shardul P. More, Tanuja S. Pawar
arXiv Machine Learning
Sep 21

Gradient-Stable Attention Heads Signal LLM Correctness

The paper introduces HeadEntropy, a training‑free method that predicts the correctness of large language model (LLM) answers by measuring how stable each attention head’s pattern is to further gradient updates. By linking the trace of the softmax Jacobian to 2‑Renyi entropy, the authors show that attention spread correlates with gradient stability, enabling accurate hallucination detection without reference annotations. Across five instruction‑tuned LLMs and five diverse datasets—including medicine, multi‑hop reasoning, and mathematics—HeadEntropy achieves a 0.736 AUROC, outperforming other training‑free baselines and matching hidden‑state probes while incurring less than 1% of inference cost.

By Sophie Ostmeier, Brian Axelrod, Maya Varma, Asad Aali, Yabin Zhang, Magdalini Paschali, Sanmi Koyejo, Curtis Langlotz, Akshay Chaudhari
arXiv Machine Learning
Sep 14

When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs

The paper examines a specific type of hallucination in large language models caused by spurious correlations—unintended, statistically prominent associations in training data such as surnames linked to nationalities. These hallucinations are confidently produced, persist regardless of model scaling or refusal fine‑tuning, and evade existing detection methods like confidence filtering and inner‑state probing. The authors use controlled synthetic experiments and evaluations on both open‑source and proprietary LLMs, including GPT‑5, to demonstrate the failure of current detection techniques and provide a theoretical explanation for why statistical biases undermine confidence‑based approaches.

By Shaowen Wang, Yiqi Dong, Ruinian Chang, Tansheng Zhu, Yuebo Sun, Kaifeng Lyu, Jian Li