arXiv AI

Why Fine-Tuning Encourages Hallucinations and How to Fix It

The paper investigates why supervised fine‑tuning (SFT) of large language models leads to increased hallucinations of factual information. It proposes a self‑distillation SFT approach that regularizes output‑distribution drift to preserve pre‑training knowledge, and shows that freezing parameter groups can reduce hallucinations when new knowledge is unnecessary. Experiments attribute the main cause to interference among overlapping semantic representations, which self‑distillation mitigates, and an associative‑memory model explains the forgetting dynamics.

arXiv AI
Aug 24

Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation

The paper identifies a specific issue in supervised fine‑tuning (SFT) of large language models called factual access failure, where models can recognize correct facts under constrained tests but fail to generate them in open‑ended settings. It demonstrates that SFT can cause both genuine wrong answers and expression‑level errors such as verbosity or formatting mismatches. To mitigate this, the authors propose Recall‑Anchored Distillation (RAD), a self‑distillation method that aligns the fine‑tuned model with the base model’s soft output distribution on unlabeled out‑of‑distribution text, thereby recovering lost factual recall without needing labeled data.

By Haodong Chen, Yadong Wang, Shengtao Wen, Dong Liang, Xiang Chen
arXiv AI
Sep 21

GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation

The paper introduces GUARD, a method for natural forgetting in large reasoning models that transforms unsafe disclosures into safe-exit trajectories using guided answer‑reasoning distillation. It aligns a frozen model with guidance tokens and distills this behavior into the parameters, aiming for a coherent, non‑disclosing chain of thought followed by a refusal‑style answer. The authors also propose the Natural Forgetting Reasoning Score (NFRS) to evaluate structural stability, fluency, and unsupported substitutes, and demonstrate GUARD’s effectiveness on R‑TOFU and a STAR‑1‑derived harmful‑intent setting.

By Zeyu Yan, Guanghao Zhou, Minghui Qiu, Ming Gao, Cen Chen
arXiv AI
Sep 21

Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration

Fact Grounded Attention (FGA) is a new architectural change that injects verifiable knowledge directly into the transformer’s pre‑softmax attention scores, turning large language models into deterministic truth‑tellers. By modifying the attention mechanism itself, FGA prevents hallucinations for facts that exist in its knowledge base, unlike prior methods that patch output or prepend retrieved text. Experiments on 1,107 technical queries show accuracy rising from 6.3% with vanilla Llama 3.2 to 99.7% with FGA, and knowledge updates can be applied in under one second without retraining.

By Aayush Gupta, Manish Choudhary
arXiv AI
Aug 11

Unified Hallucination Fuzzing for Multimodal Large Language Models

arXiv:2608. 07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications.

By Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You