arXiv:2607. 08111v1 Announce Type: cross Abstract: Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable.
By Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo, Haitao Qian, Yiming Cheng
State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which remain underrepresented in speech technology research. In this work, we introduce LuxEmo, a 21-hour conversational expressive speech corpus for Luxembourgish with 4 emotion categories.
arXiv:2606. 00012v1 Announce Type: cross Abstract: Multi-party dialogue discourse parsing aims to identify dependency structures and relation types between utterances in conversations.
By Shannan Liu, Peifeng Li, Yaxin Fan, Qiaoming Zhu
arXiv:2606. 24941v2 Announce Type: replace-cross Abstract: Reviewing recorded interviews for affective cues such as composure and agitation is slow and subjective, and cloud services that could automate the task require sensitive audio to leave the device.
By Wai Laam Mak, Isibor Kennedy Ihianle, Pedro Machado
arXiv:2609.00802v1 Announce Type: new
Abstract: Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate i...
By Taiga Mori, Koji Inoue, Mikey Elmers, Divesh Lala, Tatsuya Kawahara
arXiv:2608.20905v1 Announce Type: new
Abstract: Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue...
By Yi Zheng, Yifan Xu, Yan Zhou, Hejia Chen, Chunyu Qiang, Xiaoqiang Liu, Xiaohan Li, Shenze Huang, Yue Zhang, Guoying Zhao, Pengfei Wan
arXiv:2609.14542v1 Announce Type: new
Abstract: Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utt...
By Ahmad Amirivojdan, Farzad Nadiri, Abolfazl Alizadeh, Shaghayegh Yaraghi
TalkFa introduces a unified benchmark for Farsi dialogue generation and understanding, comprising three datasets: WIKI‑FADIAL (4.2K Wikipedia‑grounded dialogues), DAILYDIALOG‑FA (6.6K dialogues with dialogue‑act and emotion annotations), and PLAYDIAL‑FA (2.1K theatrical dialogues with sentiment labels). All dialogues are curated through multi‑stage review by native speakers, ensuring high quality. Experiments show that LoRA fine‑tuning improves generation performance with less data, while specific models excel on classification tasks, and human evaluation confirms the benchmark’s reliability.
By Neda Jamshidi, Kamyar Zeinalipour, Fahimeh Akbari, Monica Bianchini, Marco Maggini, Marco Gori
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik
MENASpeechBank is a new reference voice bank that provides about 18,000 high‑quality utterances from 124 speakers across multiple MENA countries, covering English, Modern Standard Arabic, and regional Arabic varieties. The dataset is built through a controllable pipeline that creates persona profiles inspired by the World Values Survey, defines a taxonomy of roughly 5,000 conversational scenarios, matches personas to scenarios via semantic similarity, and generates around 417,000 role‑play conversations using an LLM. Synthetic speaker‑conditioned user‑turn audio is produced from reference recordings to maintain speaker diversity, and both synthetic and human‑recorded conversations are evaluated and analyzed for quality.
By Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi, Shammur Absar Chowdhury, Firoj Alam
arXiv:2510.10774v4 Announce Type: replace-cross
Abstract: Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech...
By Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery
The paper introduces VeriSpeak, a benchmark of 3,879 spoken claims for evaluating fact verification in Large Audio Language Models (LALMs). It shows a clear modality gap: models that verify written claims well often fail on spoken versions, and retrieval alone offers limited improvement. Combining retrieval with explicit reasoning yields the best performance, reaching 86.1% accuracy and demonstrating the need for grounded reasoning over retrieved evidence in speech misinformation detection.
By Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri