arXiv AI By Lin Li, Georgia Channing, Suhaas M Bhat, Gabriel Davis Jones, Yarin Gal

Building Reliable Long-Form Generation via Hallucination Rejection Sampling

Read the original on arXiv AI →

arXiv:2606. 03628v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in open-ended text generation, yet they remain prone to hallucinating incorrect or unsupported content, which undermines their reliability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs

The paper presents a geometric framework for quantifying uncertainty in large language models (LLMs) at both the prompt and answer levels. By modeling a prompt-conditioned semantic distribution in answer embedding space and using archetypal analysis on multiple sampled answers, the method estimates distribution entropy for prompt-level uncertainty and atypicality for individual answer reliability. Experiments demonstrate comparable or superior performance to existing techniques on short-form QA datasets and notably better results on medical datasets where hallucinations pose critical risks.

By Edward Phillips, Sean Wu, Soheila Molaei, Danielle Belgrave, Anshul Thakur, David Clifton
arXiv AI
Sep 10

Evidence-Aligned Entity Verification for Hallucination Detection in Retrieval-Augmented Generation

The paper introduces Evidence-Aligned Entity Verification (EAEV), a method for detecting entity-level hallucinations in retrieval-augmented generation (RAG). EAEV aligns generated entities with retrieved evidence across three dimensions and uses counterfactual stability analysis to maintain robust alignments when evidence changes. Experiments on multiple RAG benchmarks show that EAEV consistently outperforms existing hallucination detection methods and generalizes well.

By Runsong Jia, Zhen Fang, Mengjia Wu, Jie Lu, Yi Zhang