arXiv Machine Learning

Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs

arXiv Machine Learning
Jul 14

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

arXiv:2607. 10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues.

By Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie
arXiv Computation and Language
Sep 11

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

OmniHallu is a unified framework for detecting hallucinations in multimodal large language models across both comprehension and generation tasks involving image, video, and audio modalities. It introduces OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations for six cross-modal tasks (I2T, V2T, A2T, T2I, T2V, T2A). The system uses a multi‑agent architecture that decomposes outputs into atomic claims, verifies them with modality‑specific experts, and aggregates evidence through structured reasoning, while a preference‑optimized verifier reduces expert calls by 66% with minimal performance loss.

By Jianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo
arXiv Computer Vision
Sep 17

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

arXiv:2609.18323v1 Announce Type: new Abstract: Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H...

By Haoyu Zhao, Zihao Zhao, Tianyu Deng, Ziqin Xu, Zihao Zhang, Xudong Wang, Jinxiang Guo, Chen Gao, Ziyi Ye, Yeying Jin, Jiaxi Gu, Zuxuan Wu, Shuicheng Yan
arXiv Machine Learning
Sep 23

Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation

The paper investigates why joint audio–video generators often learn to predict sound from visual appearance rather than from the underlying event, a problem termed the visual shortcut. By constructing a controlled causal model where audio is independent of video appearance, the authors show that common remedies such as shared latent spaces fail to prevent this shortcut. They propose that intervening on the nuisance appearance is necessary and sufficient for counterfactual invariance, and validate this approach across synthetic and real datasets, highlighting the remaining challenge of unknown nuisances.

By Jian Xu, Delu Zeng, John Paisley
arXiv AI
Aug 26

EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

EXAM$^2$ is a new benchmark for audio understanding that covers six languages and multiple modalities—speech, sound, music, mixed-audio, and visual images—providing 5,667 multiple-choice questions, 22,614 image instances, and 135,684 multilingual translations. It evaluates large audio language models (LALMs) and multimodal large language models (LLMs), revealing significant gaps in multilingual and cross‑modal performance. The authors also introduce Gemma3n-EXAM$^2$, a lightweight fusion model that improves multilingual results by up to 12.4% and multimodal results by 21.7% over a strong baseline.

By Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen
arXiv Computer Vision
Sep 16

Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?

Video-HolmesV2 is a new benchmark that tests multimodal large language models on their ability to reason with spatio‑temporal audio‑visual evidence in long videos. It requires models to justify answers with precise evidence, uses a multi‑model cross‑verification pipeline and a spatio‑temporal evidence‑aware metric, and introduces an audio‑text guided token compression framework to reduce long‑context noise. In evaluations, even strong proprietary models score below 60% while the proposed approach outperforms comparable open‑source omni‑models.

By Zhaoyang Wei, Zipeng Wang, Yushe Cao, Chenhui Qiang, Shuaibing Cheng, Xuesong Yang, Sen Nie, Bowen Jiang, Wenchao Ding, Yanchao Hao, Zheng Wei, Xuehui Yu, Zhenjun Han
arXiv AI
Aug 26

OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses

The paper introduces D3-Omni, a balanced and decoupled benchmark designed to diagnose fine‑grained multimodal understanding in OmniJudges that evaluate text‑to‑image, text‑to‑video, and text‑to‑speech generation. D3-Omni covers 53 orthogonal binary dimensions across 10,671 samples, using fixed positive seeds and controlled prompt rewriting to generate negatives, thereby ensuring each error can be attributed to a single capability. The benchmark’s dual‑balanced, decoupled, and dynamic design achieves near 1:1 per‑dimension parity and a uniform total‑score distribution, revealing that strong OmniJudges often miss modality‑related failures and treat distinct attributes as a single decision, masking systematic blind spots.

By Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang, Ziyi Cheng, Xinfa Zhu, HangRui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu