Hugging Face Trending Papers

Conformal Language Modeling via Posterior Sampling

Read the original on Hugging Face Trending Papers →

Large Language Models remain plagued by hallucinations. Recent work has sought to tame their prevalence using statistical techniques based on conformal prediction, with both theoretical and empirical success.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Jun 2

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

arXiv:2605. 28910v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect statements that limit their reliability in specialized healthcare applications.

By Shamanth Kuthpadi Seethakantha, Dung Ngoc Thai, Vara Prasad Gudi, Simran Tiwari, Rami Matar, Avijit Mitra, Wenlong Zhao, Andrew McCallum, Wael Salloum
arXiv AI
Aug 5

Quantifying Hallucinations in Language Language Models on Medical Textbooks

arXiv:2603. 09986v3 Announce Type: replace-cross Abstract: Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious problem within natural language processing for which we do not yet have an effective solution to mitigate against.

By Brandon C. Colelough, Davis Bartels, Dina Demner-Fushman
arXiv Computation and Language
Sep 23

Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs

The paper presents a geometric framework for quantifying uncertainty in large language models (LLMs) at both the prompt and answer levels. By modeling a prompt-conditioned semantic distribution in answer embedding space and using archetypal analysis on multiple sampled answers, the method estimates distribution entropy for prompt-level uncertainty and atypicality for individual answer reliability. Experiments demonstrate comparable or superior performance to existing techniques on short-form QA datasets and notably better results on medical datasets where hallucinations pose critical risks.

By Edward Phillips, Sean Wu, Soheila Molaei, Danielle Belgrave, Anshul Thakur, David Clifton
arXiv AI
Jun 9

CARE: A Conformal Safety Layer for Medical Summarization

arXiv:2606. 08969v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information and introduce unsupported claims.

By Suhana Bedi, Bridget Lin, Anson Y. Zhou, Chloe O. Stanwyck, Jenelle A. Jindal, Sanmi Koyejo, David Stutz, Nigam H. Shah