When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection
Read the original on arXiv Computation and Language →The study investigates how multimodal large language models (MLLMs) use prosodic cues in sarcasm detection. By testing Qwen2.5‑Omni and Qwen3‑Omni on Mandarin Chinese and English across five modality conditions, the authors find that adding audio increases false positives without improving true positives. Acoustic error analysis shows that models rely on a stereotypical prosodic pattern—elevated pitch and irregular pausing—that does not align with genuine sarcasm cues, and manipulating these dimensions alone can raise false positive rates up to 60%. The same effect appears in Gemini 3 Flash Preview, indicating the heuristic is not limited to a single architecture.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.