arXiv:2605.13737v2 Announce Type: replace
Abstract: When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in...
By Trung Nguyen Quang, Yiming Gao, Fanyi Pu, Kaichen Zhang, Shuo Sun, Ziwei Liu
arXiv:2606. 00959v1 Announce Type: new Abstract: Understanding modality interaction in multimodal large language models (MLLMs) is central to reliable deployment.
By Wanlong Fang, Tianle Zhang, Wen Tao, Alvin Chan
arXiv:2609.06011v1 Announce Type: cross
Abstract: Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexp...
By Yen-Ting Piao, Shu-Yun Chen, Chin-Hui Chu, Chun-Wei Chen, Shih-Yun Shan Kuan, Hung-yi Lee, Yun-Nung Chen
arXiv:2609.18323v1 Announce Type: new
Abstract: Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H...
By Haoyu Zhao, Zihao Zhao, Tianyu Deng, Ziqin Xu, Zihao Zhang, Xudong Wang, Jinxiang Guo, Chen Gao, Ziyi Ye, Yeying Jin, Jiaxi Gu, Zuxuan Wu, Shuicheng Yan
OmniHallu is a unified framework for detecting hallucinations in multimodal large language models across both comprehension and generation tasks involving image, video, and audio modalities. It introduces OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations for six cross-modal tasks (I2T, V2T, A2T, T2I, T2V, T2A). The system uses a multi‑agent architecture that decomposes outputs into atomic claims, verifies them with modality‑specific experts, and aggregates evidence through structured reasoning, while a preference‑optimized verifier reduces expert calls by 66% with minimal performance loss.
By Jianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo
The paper investigates how multimodal large language models (MLLMs) handle conflicting evidence presented in text, image, or both forms. Across 13 MLLMs and two datasets, the authors find that models are not robust to knowledge conflict: they tend to accept contradictory image evidence more readily than contradictory text, and when both modalities conflict the preference is arbitrary, depending on input order, model, and dataset. The instability degrades multimodal retrieval-augmented generation and can be exploited by adversarial attacks, while simple mitigation techniques such as prompting, steering, and direct preference optimization largely fail, with supervised fine‑tuning offering only moderate improvement.
By Jungyeon Lee, Yejin Yoon, Taeuk Kim