arXiv:2608.16353v2 Announce Type: replace
Abstract: Internal-state probes enable truthfulness prediction before a large language model generates an answer. When detectors change both the layers they...
By Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye, Junwei Zhang, Weiran Yao, Zhiwei Liu, Qingsong Wen, Yilei Shao
The study shows that hallucination detection in large language models is largely driven by a single mean‑shift component in hidden states. Across three 7B‑scale models and multiple datasets, removing this direction reduces detection to chance, while a simple L2‑regularized logistic regression achieves high AUROC (0.952) and outperforms more complex probe architectures. The authors introduce LayerMix, a multi‑layer aggregation method that matches oracle‑layer performance without requiring oracle access, demonstrating that apparent probe complexity stems from high‑dimensional covariance estimation rather than non‑linearity.
By Jungseob Lee, Jaehyung Seo, Heuiseok Lim
The paper investigates why vision‑language models that tokenize images with vector‑quantized (VQ) codebooks frequently hallucinate objects on grounded yes/no tasks. By applying activation patching across 25 models from eight large‑language‑model families, the authors uncover an early‑layer attention routing circuit shared by VQ‑tokenized VLMs. They develop a three‑gate diagnostic that isolates ten models carrying this circuit, show that swapping a single architectural component (VQ+Linear) introduces the circuit, and demonstrate that ablating the early‑layer ($L_0$) component reduces hallucinations in open‑ended generation by 31 % while other decoding‑time fixes do not.
By Shamanthak Hegde, Xiangrui Liu, Maitreya Patel, Yezhou Yang
Optimal transport (OT) has been shown to detect hallucinations in neural machine translation (NMT) by measuring the geometric distance between cross-attention distributions and a reference distribution, without any supervision. We extend this analysis to all six decoder layers of the Fairseq DE-EN model ($N=3{,}414$), showing that Wass-to-Unif and Wass-to-Data are complementary detectors specialised across hallucination types, that detection is concentrated in layers L1--L4 with L5 anti-predictive for subtler types, and that hallucinated translations lack the exploratory attention phase present in correct translations from the first decoding step.
The paper introduces the Memory Decision Layer (MDL), a zero‑parameter controller that sits between retrieval and generation in large language models. MDL uses a three‑signal complementary encoder—combining relevance, reliability, and task risk—to produce an interpretable decision about the trustworthiness of retrieved memories. By decoupling confidence from consistency and enabling risk inversion and abstention, MDL cuts hallucination rates under conflicting memories by roughly 56% and nearly eliminates them in high‑risk scenarios, all while adding only 0.14 ms per decision.
By Yiming Zhang, Jinghong Zhang, Haoran Zhao, Yiren Ma, Chunlei Zhao
arXiv:2606. 19404v1 Announce Type: new Abstract: Hallucination detection in large language models (LLMs) is deployment-critical, and recent work shows that the spectrum of attention-derived graph Laplacians carries strong signal about reasoning quality.
By Salim Khazem
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2607. 02893v1 Announce Type: new Abstract: Low-bit quantization shrinks language models but treats precision as a single global hyper-parameter: every weight uses the same bit-width.
By Hamish Ogilvy
The paper introduces a multi‑signal pipeline for detecting hallucinations in large language models, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves high performance (F1 = 0.915, AUROC = 0.977) across QA, summarization, and dialogue, and shows that 25 % of training data yields 77 % of full‑data performance. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator cuts hallucination rates from 85.5 % to 37.7 %, and that domain‑specific fine‑tuning (PubMedBERT on SciFact) outperforms general‑domain models for biomedical text.
By Varun Teja Chundru, Debasmita Biswas
arXiv:2608. 16353v1 Announce Type: cross Abstract: Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments.
By Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye, Junwei Zhang, Weiran Yao, Zhiwei Liu, Qingsong Wen, Yilei Shao
The paper introduces MACCHIATO, a training algorithm that builds a ReLU‑MLP from partial truth‑table data while simultaneously constructing an explicit Boolean circuit over AND, OR, and XOR gates that certifies the network’s computation. The method iteratively projects residuals onto low‑dimensional Boolean classes, compiles the resulting circuit into a ReLU‑MLP, and uses logic minimization and influence‑based variable selection to achieve a six‑layer network with provable truth‑table error bounds. Experiments on synthetic random‑junta tasks show that these certified networks outperform Adam‑trained MLPs in data‑sparse or projection‑aligned regimes and complete faster than flat ESPRESSO in certain settings.
By Hrad Ghoukasian, Anastasis Kratsios
The study investigates how post‑training quantization (PTQ) affects proactive interference (PI) in large language models. Using bitsandbytes, the authors compare FP16, INT8, and INT4/NF4 precision across three instruction‑tuned models and find that INT4 quantization markedly degrades accuracy under high interference, with INT8 also incurring a smaller penalty in two of the three models. The degradation is linked to increased same‑key intrusion errors and originates in the quantized transformer backbone rather than the output layer.
By Shayan Shahrabi-Farahani (Shahid Beheshti University, Tehran, Iran), Dara Rahmati (Shahid Beheshti University, Tehran, Iran)