arXiv AI By Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu, Tae-Hyun Oh

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

Read the original on arXiv AI →

arXiv:2606. 02642v1 Announce Type: cross Abstract: Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.