arXiv Machine Learning By Aizierjiang Aiersilan

Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs

Read the original on arXiv Machine Learning →

arXiv:2606. 02628v1 Announce Type: new Abstract: We investigate whether open-source LLMs encode a linearly separable truthfulness signal in their hidden states, and at which network depth this signal is strongest.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 1

The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

The study shows that hallucination detection in large language models is largely driven by a single mean‑shift component in hidden states. Across three 7B‑scale models and multiple datasets, removing this direction reduces detection to chance, while a simple L2‑regularized logistic regression achieves high AUROC (0.952) and outperforms more complex probe architectures. The authors introduce LayerMix, a multi‑layer aggregation method that matches oracle‑layer performance without requiring oracle access, demonstrating that apparent probe complexity stems from high‑dimensional covariance estimation rather than non‑linearity.

By Jungseob Lee, Jaehyung Seo, Heuiseok Lim
arXiv Machine Learning
Sep 25

Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models

The paper investigates why vision‑language models that tokenize images with vector‑quantized (VQ) codebooks frequently hallucinate objects on grounded yes/no tasks. By applying activation patching across 25 models from eight large‑language‑model families, the authors uncover an early‑layer attention routing circuit shared by VQ‑tokenized VLMs. They develop a three‑gate diagnostic that isolates ten models carrying this circuit, show that swapping a single architectural component (VQ+Linear) introduces the circuit, and demonstrate that ablating the early‑layer ($L_0$) component reduces hallucinations in open‑ended generation by 31 % while other decoding‑time fixes do not.

By Shamanthak Hegde, Xiangrui Liu, Maitreya Patel, Yezhou Yang
Hugging Face Trending Papers
Jun 11

Layer-Resolved Optimal Transport for Hallucination Detection in NMT and Abstractive Summarization

Optimal transport (OT) has been shown to detect hallucinations in neural machine translation (NMT) by measuring the geometric distance between cross-attention distributions and a reference distribution, without any supervision. We extend this analysis to all six decoder layers of the Fairseq DE-EN model ($N=3{,}414$), showing that Wass-to-Unif and Wass-to-Data are complementary detectors specialised across hallucination types, that detection is concentrated in layers L1--L4 with L5 anti-predictive for subtler types, and that hallucinated translations lack the exploratory attention phase present in correct translations from the first decoding step.

arXiv Computation and Language
Sep 21

An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

The paper introduces the Memory Decision Layer (MDL), a zero‑parameter controller that sits between retrieval and generation in large language models. MDL uses a three‑signal complementary encoder—combining relevance, reliability, and task risk—to produce an interpretable decision about the trustworthiness of retrieved memories. By decoupling confidence from consistency and enabling risk inversion and abstention, MDL cuts hallucination rates under conflicting memories by roughly 56% and nearly eliminates them in high‑risk scenarios, all while adding only 0.14 ms per decision.

By Yiming Zhang, Jinghong Zhang, Haoran Zhao, Yiren Ma, Chunlei Zhao