arXiv AI

ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning

arXiv:2606. 14697v1 Announce Type: cross Abstract: Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision support.

arXiv AI
Jul 29

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

arXiv:2607. 25933v1 Announce Type: cross Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning.

By Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, Nicol\'as Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shijian Lu, Douglas Teodoro, Naoto Yokoya, Ross Koppel, Mona Diab, Hua Xu, David W. Bates, Nan Liu, Yifan Peng
arXiv AI
Jul 21

Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain

arXiv:2603. 21693v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have shown strong potential for medical Visual Question Answering (VQA), yet they remain prone to hallucinations, defined as generating responses that contradict the input image, posing serious risks in clinical settings.

By Mohammad Asadi, Tahoura Nedaee, Jack W. O'Sullivan, Euan Ashley, Ehsan Adeli
arXiv AI
Sep 7

Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection

The paper explores using a language model’s low‑level symbolic skill—specifically SQL—to detect hallucinations without fine‑tuning. By having the model construct an SQL database from reference documents, it can reason over both the source and the model’s output, creating a neurosymbolic check. Experiments on RAGTruth and DiaHalu show this method outperforms direct prediction and rivals state‑of‑the‑art detectors, highlighting the value of leveraging inherent symbolic competences in LLMs.

By Renato Vukovic, Hsien-chin Lin, Carel van Niekerk, Benjamin Ruppik, Michael Heck, Shutong Feng, Nurul Lubis, Milica Gasic
arXiv AI
Sep 15

Hallucination in Multimodal Foundation Models: A Survey on Causes, Corrections, and Evaluations

The article surveys hallucination issues in Large Vision‑Language Models (LVLMs), a type of multimodal foundation model that blends visual data with large language models. It categorizes hallucination causes into model architecture and data quality, presents a taxonomy of mitigation strategies, and critically evaluates existing evaluation benchmarks from both discriminative and generative viewpoints. The survey also outlines open challenges and future research directions to improve LVLM reliability and trustworthiness.

By Yinghao Guo, Wei Lan, Wenyi Chen, Qingfeng Chen, Shichao Zhang, Shirui Pan, Huiyu Zhou, Yi Pan
arXiv AI
Aug 5

UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space

arXiv:2608. 03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence.

By Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad Samimi, Mohammad Hossein Rohban
arXiv AI
Aug 18

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.

By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang