arXiv:2606. 07533v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) effectively integrate text and audio to interpret context in complex interactive dialogues.
By Pawe{\l} Pozorski, Jakub Muszy\'nski, Maria Ganzha
This paper investigates whether prosodic features—pitch, energy, and timing—are preserved when speech is translated between languages. Using multilingual dubbing data for English‑German, English‑Spanish, and English‑French pairs, the authors conduct a fine‑grained cross‑lingual analysis to quantify similarities and differences in prosody. The study identifies inherent cross‑lingual correlations in prosodic structure and explores how linguistic and alignment factors influence these patterns.
By Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman, Philipp Koehn
The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.
By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv:2609.16396v1 Announce Type: new
Abstract: Negation is typically modeled through its linguistic realization, although spoken interaction is accompanied by tightly coordinated nonverbal behavior....
By Leon Hammerla, Patrick Schrottenbacher, Alexander Mehler
The study investigates how multimodal large language models (MLLMs) use prosodic cues in sarcasm detection. By testing Qwen2.5‑Omni and Qwen3‑Omni on Mandarin Chinese and English across five modality conditions, the authors find that adding audio increases false positives without improving true positives. Acoustic error analysis shows that models rely on a stereotypical prosodic pattern—elevated pitch and irregular pausing—that does not align with genuine sarcasm cues, and manipulating these dimensions alone can raise false positive rates up to 60%. The same effect appears in Gemini 3 Flash Preview, indicating the heuristic is not limited to a single architecture.
By Yongjian Chen, Pengfei Wei, Yiqun Sun, Zhu Li, Lawrence B. Hsieh
Multimodal Large Language Models (MLLMs) process speech and text jointly, yet whether they exploit prosodic cues for pragmatic inference or rely on surface acoustic patterns has received little system...
arXiv:2603.23938v2 Announce Type: replace
Abstract: Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak...
By Seunghee Kim, Bumkyu Park, Kyudan Jung, Joosung Lee, Soyoon Kim, Jeonghoon Kim, Taeuk Kim, Hwiyeol Jo
arXiv:2609.14743v1 Announce Type: new
Abstract: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify...
By Ju-Chieh Chou, Jiawei Zhou, Karen Livescu
The paper examines whether adding syntactic and rhetorical structure to text can improve the prediction of incoherence in large language model outputs. Experiments show that plain text actually yields higher accuracy, as the added structural information conflicts with the models’ architectures. The authors also demonstrate that coherence assessment can help detect misleading content by applying zero‑shot experiments to a Brazilian disinformation dataset.
By Victor Mazzotti, Luiz Pereira, Marina Bitencourt dos Santos, Helena Maia, Carlos Caetano, N\'adia Felix, Sandra Avila
The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.
By Titus von der Malsburg, Sebastian Pad\'o
arXiv:2607. 06611v1 Announce Type: cross Abstract: Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and the interpretation of uttered words.
By Andrei-George Durdun, Victor Constantinescu, Radu Tudor Ionescu
arXiv:2606. 11371v1 Announce Type: cross Abstract: Spoken language, whether produced by humans or large language models (LLM), unfolds over time with varying semantic content.
By Han-Jen Chang, Yasir \c{C}atal, Angelika Wolman, Agust\'in Ib\'a\~nez, David Smith, I-Wen Su, Kai-Yuan Cheng, Georg Northoff