CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The study investigates how multimodal large language models (MLLMs) use prosodic cues in sarcasm detection. By testing Qwen2.5‑Omni and Qwen3‑Omni on Mandarin Chinese and English across five modality conditions, the authors find that adding audio increases false positives without improving true positives. Acoustic error analysis shows that models rely on a stereotypical prosodic pattern—elevated pitch and irregular pausing—that does not align with genuine sarcasm cues, and manipulating these dimensions alone can raise false positive rates up to 60%. The same effect appears in Gemini 3 Flash Preview, indicating the heuristic is not limited to a single architecture.
Multimodal Large Language Models (MLLMs) process speech and text jointly, yet whether they exploit prosodic cues for pragmatic inference or rely on surface acoustic patterns has received little system...
arXiv:2606. 06168v1 Announce Type: new Abstract: We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local prosodic dynamics and the utterance-level emotional baseline.
Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions.
arXiv:2608. 09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.
arXiv:2607. 06611v1 Announce Type: cross Abstract: Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and the interpretation of uttered words.