arXiv AI

NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass

arXiv AI
Jun 9

BEACON: Behavioral Entropy Aggregation for Cross-Model Hallucination Detection in Large Language Models

arXiv:2606. 07528v1 Announce Type: cross Abstract: Hallucination in large language models (LLMs), defined as the generation of factually incorrect or unsupported content, remains a critical barrier to reliable deployment.

By Naveen Bera, Pulijala Sai Nikhila, Kondaguduru Abhiram, Shaik Gayaz Ali, Shoaib Sadiq Salehmohamed, Shaik Mohammed Omar, Jinal Prashant Thakkar, Hansika Aredla, Shalmali Ayachit
arXiv AI
4d ago

MedHal: a Synthetic Dataset for Medical Hallucination Detection

MedHal is a large-scale synthetic dataset created to detect hallucinations in medical AI-generated text. It includes diverse medical sources and tasks that cover both intrinsic and extrinsic hallucinations, providing a substantial volume of samples for training. The authors demonstrate that models trained on MedHal outperform general-purpose hallucination detectors, highlighting its usefulness for medical AI development.

By Fabrice Lamarche, Gaya Mehenni, Neshat Elhami Fard, Odette Rios-Ibacache, Li Ming Wang, John Kildea, Amal Zouaq
Hugging Face Trending Papers
Aug 18

Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

The paper introduces InnerExpert, a method that uses Mixture-of-Experts (MoE) internal signals—such as router entropy, expert disagreement, and usage patterns—to detect hallucinations at the token level in large language models. By combining these MoE-specific signals with standard transformer features into compact per-token vectors, InnerExpert trains a lightweight detector using an LLM-as-a-judge pipeline, enabling continuous updates without manual labeling. Experiments across five datasets and two MoE architectures show that InnerExpert outperforms existing methods, achieving up to 0.91 answer-level and 0.76 token-level AUROC with only a single forward pass.