arXiv:2609.15128v1 Announce Type: new
Abstract: Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support...
By Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang, Difan Zou
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, re...
OmniHallu is a unified framework for detecting hallucinations in multimodal large language models across both comprehension and generation tasks involving image, video, and audio modalities. It introduces OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations for six cross-modal tasks (I2T, V2T, A2T, T2I, T2V, T2A). The system uses a multi‑agent architecture that decomposes outputs into atomic claims, verifies them with modality‑specific experts, and aggregates evidence through structured reasoning, while a preference‑optimized verifier reduces expert calls by 66% with minimal performance loss.
By Jianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo
arXiv:2609.37568v1 Announce Type: new
Abstract: Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual,...
By Yu Zhang, Pingrui Zhang, Xuefeng Bai, Pengfei Zhang, Yang Xiang, Kehai Chen
The paper introduces IMAVB, a 500‑clip benchmark that tests whether omnimodal large language models can detect when a textual premise contradicts their visual or audio input. Experiments on eight open‑source models and Gemini 3.1 Pro reveal a Representation‑Action Gap: internal states encode mismatches, yet the models rarely reject false premises, exhibiting under‑rejection or over‑rejection. A probe‑guided logit adjustment improves rejection behavior, suggesting the main bottleneck is in translating perception to action rather than in perception itself.
By Trung Nguyen Quang, Yiming Gao, Fanyi Pu, Kaichen Zhang, Shuo Sun, Ziwei Liu
The paper examines six inference-time hallucination mitigation methods applied to three large vision-language models across four benchmarks, including MMStar. It finds that reducing hallucination rates often comes at the cost of lower informativeness—such as decreased object recall, visual coverage, and response detail—and that gains on hallucination benchmarks do not consistently translate to improved performance on fine-grained perception and reasoning tasks. The authors argue that current evaluation protocols may overstate progress by favoring conservative generation, and propose that hallucination mitigation should be assessed as a trade-off among faithfulness, informativeness, and overall capability.
By Mehrdad Fazli, Sina Mansouri, Mohit Marvania, Ziwei Zhu