arXiv:2603.23938v2 Announce Type: replace
Abstract: Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak...
By Seunghee Kim, Bumkyu Park, Kyudan Jung, Joosung Lee, Soyoon Kim, Jeonghoon Kim, Taeuk Kim, Hwiyeol Jo
arXiv:2608. 10720v1 Announce Type: new Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied.
By Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo
OmniHallu is a unified framework for detecting hallucinations in multimodal large language models across both comprehension and generation tasks involving image, video, and audio modalities. It introduces OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations for six cross-modal tasks (I2T, V2T, A2T, T2I, T2V, T2A). The system uses a multi‑agent architecture that decomposes outputs into atomic claims, verifies them with modality‑specific experts, and aggregates evidence through structured reasoning, while a preference‑optimized verifier reduces expert calls by 66% with minimal performance loss.
By Jianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo
arXiv:2505.17613v2 Announce Type: replace
Abstract: Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation...
By Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu
The Modality Maturity Index (MMI) is a new benchmark that evaluates large language models on their ability to handle five different modalities—text, image, audio, video, and document—across up to three-input and three-output combinations. It contains 893 self‑contained questions, each with human‑authored rubric criteria for the expected output modalities, and measures performance via an MMI Value and a Modality Presence Score (MPS). Experiments on five frontier multimodal models show low MPS scores, indicating limited modality availability, and confirm that LLM judges can reliably assess output correctness against human‑blind rubric scoring on 70.8% of cases.
By Rohit Patel, Dieuwke Hupkes, Sloan Strader
The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.
By Ankan Deria, Hanoona Rasheed, Xilin He, Fahad Shahbaz Khan, Salman Khan
arXiv:2607. 02963v1 Announce Type: cross Abstract: Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation.
By Wenzheng Zeng, Siyi Jiao, Chen Gao, Hwee Tou Ng, Mike Zheng Shou
arXiv:2604. 25220v2 Announce Type: replace Abstract: Data videos combine animated visualizations with synchronized narration to communicate quantitative information and are widely used in journalism, education, and public communication.
By Ridwan Mahbub, Syem Aziz, Mizanur Rahman, Mahir Ahmed, Shadikur Rahman, Shafiq Joty, Enamul Hoque
OmniFusion is an end‑to‑end multilingual multimodal translation system that fuses a pretrained multimodal foundation model (Omni 2.5‑7B) with a translation large language model (SeedX PPO‑7B). By connecting hidden states from multiple layers of the multimodal model to the translation LLM, OmniFusion can translate speech, speech‑and‑image, and text‑and‑image inputs while reducing simultaneous speech‑translation latency by about one second compared to cascaded pipelines. The approach improves overall translation quality by leveraging both audio and visual context.
By Sai Koneru, Matthias Huck, Jan Niehues
The paper presents a new approach to automatic audio description (AD) that treats the task as a constrained global optimization problem. It jointly decides what visual content is narratively important, when it can be spoken without overlapping dialogue, and how to phrase it within time limits. Using large language models for salience estimation and a mixed‑integer linear program for scheduling, the system outperforms prior methods on the REFRAMED benchmark, especially in temporal placement and narrative relevance.
By Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller
The paper introduces Omni-Interactive Universal Embedder (OmniUE), a unified embedding framework that learns a single representation space for text, video, and audio using learnable tokens and intermediate-layer representations. OmniUE supports omni-interactive querying, allowing users to input text, visual regions, or audio spans, which are processed by segmenters and an omni-LLM to generate user-conditioned embeddings. The authors evaluate OmniUE on the new OmniCHOIR benchmark and other multimodal tasks, reporting significant performance gains over state‑of‑the‑art baselines across textual, audio, and visual interactive settings.
By Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji
The paper introduces D3-Omni, a balanced and decoupled benchmark designed to diagnose fine‑grained multimodal understanding in OmniJudges that evaluate text‑to‑image, text‑to‑video, and text‑to‑speech generation. D3-Omni covers 53 orthogonal binary dimensions across 10,671 samples, using fixed positive seeds and controlled prompt rewriting to generate negatives, thereby ensuring each error can be attributed to a single capability. The benchmark’s dual‑balanced, decoupled, and dynamic design achieves near 1:1 per‑dimension parity and a uniform total‑score distribution, revealing that strong OmniJudges often miss modality‑related failures and treat distinct attributes as a single decision, masking systematic blind spots.
By Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang, Ziyi Cheng, Xinfa Zhu, HangRui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu