arXiv Machine Learning

Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

arXiv:2607. 07993v1 Announce Type: cross Abstract: Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data.

arXiv AI
2d ago

External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

The paper proposes a hidden‑state probing method for detecting hallucinations at the span level in large language model outputs, moving beyond token‑wise binary classification. By examining layer‑wise activation patterns, the approach identifies the exact onset and continuation tokens of hallucinations, achieving higher precision‑recall AUC than random baselines despite class imbalance. Additionally, the authors introduce a cross‑model detection framework where one model observes another’s internal representations, showing that an external observer can match or surpass the generator’s own self‑detection of hallucination onsets, even when the observer is smaller.

By Kingshuk Gupta, Davide Buscaldi
Hugging Face Trending Papers
Jun 25

Hallucination in World Models is Predictable and Preventable

Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics. We hypothesize that hallucination concentrates in low-coverage regions of the state-action space, where lightweight data-centric signals can both detect it and guide mitigation.

arXiv Machine Learning
Aug 12

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

arXiv:2608. 10835v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by the visual input.

By Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca, Ethan Fetaya, Yftah Ziser, Gal Chechik, Haggai Maron
arXiv Computer Vision
Sep 3

Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods

The paper examines six inference-time hallucination mitigation methods applied to three large vision-language models across four benchmarks, including MMStar. It finds that reducing hallucination rates often comes at the cost of lower informativeness—such as decreased object recall, visual coverage, and response detail—and that gains on hallucination benchmarks do not consistently translate to improved performance on fine-grained perception and reasoning tasks. The authors argue that current evaluation protocols may overstate progress by favoring conservative generation, and propose that hallucination mitigation should be assessed as a trade-off among faithfulness, informativeness, and overall capability.

By Mehrdad Fazli, Sina Mansouri, Mohit Marvania, Ziwei Zhu