What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2603.23938v2 Announce Type: replace Abstract: Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak...
arXiv:2608. 10720v1 Announce Type: new Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied.
OmniHallu is a unified framework for detecting hallucinations in multimodal large language models across both comprehension and generation tasks involving image, video, and audio modalities. It introduces OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations for six cross-modal tasks (I2T, V2T, A2T, T2I, T2V, T2A). The system uses a multi‑agent architecture that decomposes outputs into atomic claims, verifies them with modality‑specific experts, and aggregates evidence through structured reasoning, while a preference‑optimized verifier reduces expert calls by 66% with minimal performance loss.
arXiv:2505.17613v2 Announce Type: replace Abstract: Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation...
The Modality Maturity Index (MMI) is a new benchmark that evaluates large language models on their ability to handle five different modalities—text, image, audio, video, and document—across up to three-input and three-output combinations. It contains 893 self‑contained questions, each with human‑authored rubric criteria for the expected output modalities, and measures performance via an MMI Value and a Modality Presence Score (MPS). Experiments on five frontier multimodal models show low MPS scores, indicating limited modality availability, and confirm that LLM judges can reliably assess output correctness against human‑blind rubric scoring on 70.8% of cases.
The paper introduces Training-Free Omni (TFO), a plug‑and‑play framework that transforms a frozen vision‑language model (VLM) into a speech‑centric omni model without modifying its architecture or requiring multimodal re‑alignment. TFO leverages Whisper to generate confidence‑filtered, timestamped transcripts and routes them through the VLM’s existing language interface, leaving the visual pathway untouched. Evaluations on 56 benchmarks across 21 languages show that TFO matches or surpasses native omni models on audio‑visual tasks, improves audio‑only performance, and preserves strong visual and reasoning capabilities.