arXiv Computation and Language By Yangfan Hu, Xuhan Tong, Haoyue Bai, Xi Ding, Shashank Muralidhar Bharadwaj, Siyang Cao, Robert Nowak, Jiawei Zhang

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

Read the original on arXiv Computation and Language →

The paper investigates why large language models hallucinate by proposing that failures often stem from inference misalignment rather than missing knowledge. It introduces a latent key-task model that shows pretraining-frequency imbalance can cause shortcut inference paths to dominate, leading to hallucinations. The authors create TrapQA, a diagnostic testbed with ScientistQA and Real-Life Constrained QA, to demonstrate two failure modes—task-retrieval bias and key-selection bias—where biased latent inference produces hallucinated answers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 10

Evidence-Aligned Entity Verification for Hallucination Detection in Retrieval-Augmented Generation

The paper introduces Evidence-Aligned Entity Verification (EAEV), a method for detecting entity-level hallucinations in retrieval-augmented generation (RAG). EAEV aligns generated entities with retrieved evidence across three dimensions and uses counterfactual stability analysis to maintain robust alignments when evidence changes. Experiments on multiple RAG benchmarks show that EAEV consistently outperforms existing hallucination detection methods and generalizes well.

By Runsong Jia, Zhen Fang, Mengjia Wu, Jie Lu, Yi Zhang
arXiv Machine Learning
Sep 17

Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models

The paper proposes using the temporal volatility of internal attention mechanisms—measured by an unsupervised attention dispersion metric—as a diagnostic signal for hallucinations in large language models. It demonstrates that spikes in attention entropy within intermediate layers correlate with reasoning breakdowns, and shows statistically significant AUC improvements of up to +0.076 over output-based baselines on GSM8K and MATH-500 benchmarks using the Qwen2.5 model family.

By Shardul P. More, Tanuja S. Pawar