arXiv Computation and Language By Basel Mousi, Fahim Dalvi, Shammur Chowdhury, Firoj Alam, Nadir Durrani

Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models

Read the original on arXiv Computation and Language →

The paper introduces a new benchmark called Speech-Augmented Visually Grounded Contrastive Triplet Benchmark, comprising 10,150 images from 18 MENA countries, each paired with a supported statement and two plausible but unsupported alternatives. It defines contrastive instability as the rate at which multimodal models fail to resolve all statements within a triplet, distinguishing fragmented reasoning from complete failure. Experiments on recent multimodal models show that shifts in modality (text vs. speech) and language (English vs. Arabic) lead to significant triplet-level inconsistencies, especially when speech is used, which are not fully reflected by overall accuracy metrics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 26

EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

EXAM$^2$ is a new benchmark for audio understanding that covers six languages and multiple modalities—speech, sound, music, mixed-audio, and visual images—providing 5,667 multiple-choice questions, 22,614 image instances, and 135,684 multilingual translations. It evaluates large audio language models (LALMs) and multimodal large language models (LLMs), revealing significant gaps in multilingual and cross‑modal performance. The authors also introduce Gemma3n-EXAM$^2$, a lightweight fusion model that improves multilingual results by up to 12.4% and multimodal results by 21.7% over a strong baseline.

By Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen