Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models
Read the original on arXiv Computation and Language →The paper introduces a new benchmark called Speech-Augmented Visually Grounded Contrastive Triplet Benchmark, comprising 10,150 images from 18 MENA countries, each paired with a supported statement and two plausible but unsupported alternatives. It defines contrastive instability as the rate at which multimodal models fail to resolve all statements within a triplet, distinguishing fragmented reasoning from complete failure. Experiments on recent multimodal models show that shifts in modality (text vs. speech) and language (English vs. Arabic) lead to significant triplet-level inconsistencies, especially when speech is used, which are not fully reflected by overall accuracy metrics.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.