When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
Read the original on arXiv Machine Learning →The paper examines a specific type of hallucination in large language models caused by spurious correlations—unintended, statistically prominent associations in training data such as surnames linked to nationalities. These hallucinations are confidently produced, persist regardless of model scaling or refusal fine‑tuning, and evade existing detection methods like confidence filtering and inner‑state probing. The authors use controlled synthetic experiments and evaluations on both open‑source and proprietary LLMs, including GPT‑5, to demonstrate the failure of current detection techniques and provide a theoretical explanation for why statistical biases undermine confidence‑based approaches.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.