arXiv:2606. 27717v1 Announce Type: cross Abstract: Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech.
By Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang, Haonan Chen, Zeyu Jin
arXiv:2605.28227v2 Announce Type: replace
Abstract: Speech translation models are increasingly capable of preserving speech-specific information (e.g., speaker gender, prosody, and emphasis), yet eva...
By Maike Z\"ufle, Danni Liu, Vil\'em Zouhar, Jan Niehues
The Public Discourse Corpus (PDC) is the first dataset of public‑figure interview speech annotated for affective valence and epistemic modality. It contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words). A key methodological contribution is Target Speaker Participation (TSP), a five‑category annotation taxonomy with documented inter‑annotator reliability (κ = 0.616), and an audio‑first diarization pipeline that separates target‑speaker turns from interviewer and third‑party speech. The corpus, annotation tools, validation sample, and processing pipeline are released as open source.
By Bo Chen
The paper introduces ACERT, a module that incorporates a flexible-length window of conversational context to enhance Speech Emotion Recognition (SER). By capturing emotional evolution across utterances, ACERT outperforms state‑of‑the‑art methods on IEMOCAP, sets a new context‑aware benchmark on SAFE, and achieves strong results on MELD. Ablation studies attribute ACERT’s improvements to emotional and conversational continuity rather than speaker identity or acoustic conditions.
By Arthur Peuvot, Romaric Besan\c{c}on, Ga\"el de Chalendar, Bianca Vieru, Ioana Vasilescu
This paper investigates whether prosodic features—pitch, energy, and timing—are preserved when speech is translated between languages. Using multilingual dubbing data for English‑German, English‑Spanish, and English‑French pairs, the authors conduct a fine‑grained cross‑lingual analysis to quantify similarities and differences in prosody. The study identifies inherent cross‑lingual correlations in prosodic structure and explores how linguistic and alignment factors influence these patterns.
By Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman, Philipp Koehn
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leavi...
arXiv:2609.14231v1 Announce Type: cross
Abstract: Controllable synthesis of nonverbal vocalizations (NVVs) is es- sential for natural and expressive speech, but remains challeng- ing due to their aco...
By Ziyu Zhang, Yun Chen, Taihui Wang, Hanzhao Li, Qicong Xie, Rilin Chen, Zhixian Zhao, Lei Xie
arXiv:2609.14743v1 Announce Type: new
Abstract: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify...
By Ju-Chieh Chou, Jiawei Zhou, Karen Livescu
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik
arXiv:2606. 24941v2 Announce Type: replace-cross Abstract: Reviewing recorded interviews for affective cues such as composure and agitation is slow and subjective, and cloud services that could automate the task require sensitive audio to leave the device.
By Wai Laam Mak, Isibor Kennedy Ihianle, Pedro Machado
arXiv:2608. 02235v1 Announce Type: cross Abstract: Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages.
By Ali Jafar, Amal Sarmad, Shifa Yousaf, Maryam Bashir