Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2608.30270v1 Announce Type: new Abstract: Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-base...
arXiv:2507. 19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework.
The paper introduces a new benchmark called Speech-Augmented Visually Grounded Contrastive Triplet Benchmark, comprising 10,150 images from 18 MENA countries, each paired with a supported statement and two plausible but unsupported alternatives. It defines contrastive instability as the rate at which multimodal models fail to resolve all statements within a triplet, distinguishing fragmented reasoning from complete failure. Experiments on recent multimodal models show that shifts in modality (text vs. speech) and language (English vs. Arabic) lead to significant triplet-level inconsistencies, especially when speech is used, which are not fully reflected by overall accuracy metrics.
EXAM$^2$ is a new benchmark for audio understanding that covers six languages and multiple modalities—speech, sound, music, mixed-audio, and visual images—providing 5,667 multiple-choice questions, 22,614 image instances, and 135,684 multilingual translations. It evaluates large audio language models (LALMs) and multimodal large language models (LLMs), revealing significant gaps in multilingual and cross‑modal performance. The authors also introduce Gemma3n-EXAM$^2$, a lightweight fusion model that improves multilingual results by up to 12.4% and multimodal results by 21.7% over a strong baseline.
arXiv:2605. 18160v2 Announce Type: replace-cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information.
arXiv:2608. 11907v2 Announce Type: replace-cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.