arXiv:2609.15106v1 Announce Type: new
Abstract: Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a laten...
By Xuhan Tong, Jiawei Zhang
arXiv:2603. 01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisions through verbalized reasoning.
By Kyle Cox, Darius Kianersi, Adri\`a Garriga-Alonso
The paper introduces Evidence-Aligned Entity Verification (EAEV), a method for detecting entity-level hallucinations in retrieval-augmented generation (RAG). EAEV aligns generated entities with retrieved evidence across three dimensions and uses counterfactual stability analysis to maintain robust alignments when evidence changes. Experiments on multiple RAG benchmarks show that EAEV consistently outperforms existing hallucination detection methods and generalizes well.
By Runsong Jia, Zhen Fang, Mengjia Wu, Jie Lu, Yi Zhang
The paper proposes using the temporal volatility of internal attention mechanisms—measured by an unsupervised attention dispersion metric—as a diagnostic signal for hallucinations in large language models. It demonstrates that spikes in attention entropy within intermediate layers correlate with reasoning breakdowns, and shows statistically significant AUC improvements of up to +0.076 over output-based baselines on GSM8K and MATH-500 benchmarks using the Qwen2.5 model family.
By Shardul P. More, Tanuja S. Pawar
arXiv:2610.08026v1 Announce Type: new
Abstract: In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked w...
By Jorma Valjakka, Juhani Kivim\"aki, Juha Myll\"ari, Jukka K. Nurminen
arXiv:2602.11166v2 Announce Type: replace-cross
Abstract: Parameter-efficient fine-tuning (PEFT) methods are widely used to adapt large language models (LLMs) to downstream tasks and are often assume...
By Xu Hu, Yifan Zhang, Songtao Wei, Chen Zhao, Qiannan Li, Bingzhe Li, Feng Chen
arXiv:2601.08058v2 Announce Type: replace-cross
Abstract: Chain-of-Thought (CoT) prompting often improves the reasoning performance of large language models (LLMs), but the internal signal that trigg...
By Zhenghao He, Guangzhi Xiong, Bohan Liu, Sanchit Sinha, Aidong Zhang
arXiv:2606. 07537v1 Announce Type: cross Abstract: Large language models hallucinate--producing fluent, confident, factually wrong outputs--with a consistency that persists across generations and scales.
By Md. Rejaul Korim Sadi, Toufiqur Rahman Tasin, Golam Mostofa Naeem
arXiv:2605. 26795v2 Announce Type: replace Abstract: Chain-of-thought (CoT) prompting enhances large language model performance, yet what drives these gains remains unclear.
By Xiang Wang, Wei Wei
The paper investigates how adjectival modifiers affect the semantic plausibility of events, using the Adept benchmark of 16,000 English sentence pairs that differ by a single adjective. Experiments show that sentence transformers, despite being conceptually suited to the task, underperform compared to models like RoBERTa. The authors provide an error analysis and discuss the implications of their findings for future work on balancing training and test data.
By Anna Golub, Beate Zywietz, Annerose Eichel
The paper introduces TRACE, a fine‑tuning framework for Retrieval‑Augmented Generation (RAG) that addresses conflicts between retrieved knowledge and a model’s internal knowledge. TRACE uses multi‑agent debate traces to identify correct and incorrect candidates and answer‑shift patterns, providing fine‑grained supervision for reliable knowledge‑source selection. It also incorporates an answer‑completeness regularization mechanism to prevent empty, overly short, or prematurely terminated responses, thereby improving robustness against misleading retrieved content and enhancing answer quality.
By Zhengchen Huang, Yundong Sun, Minrui Song, Shuanglong Yao, Ye Liu, Ji Chen, Xing Wang
arXiv:2609.24238v1 Announce Type: new
Abstract: We reproduce and stress-test the work of Yu et al. (2023), who characterize how language models (LMs) arbitrate between memorized knowledge and contrad...
By Guilhem Fouilh\'e, Nicholas Asher, Philippe Muller