arXiv:2607. 05365v1 Announce Type: cross Abstract: Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech.
By Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar, Ashish Hallur, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez
arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.
By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
arXiv:2607. 22658v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited.
By Yuzhe Wang (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Thomas Thebaud (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Jennifer Hu (Department of Cognitive Science, Johns Hopkins University, Baltimore, USA), Jes\'us Villalba-Lopez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Venkatesh Ravichandran (Amazon AGI, USA), Georgi Tinchev (Amazon Research, UK), Najim Dehak (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Laureano Moro-Vel\'azquez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA)
arXiv:2507.18061v4 Announce Type: replace-cross
Abstract: Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks pri...
By Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang, Jing Zhou, Hang Lv, Mengjie Du, Yaodong Song, Jie Lian, Jian Kang, Jie Li, Yongxiang Li
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik
The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.
By Ilija Subasic, Andrew Rabinovich, Zhao Chen
arXiv:2609.22452v1 Announce Type: new
Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient in...
By Xize Cheng, Wenxu Jia, Chenyuhao Wen, Dongjie Fu, Zehan Wang, Xinyu Zhang, Tao Jin
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.
The paper introduces a method called randomized intermediate guidance for training tandem speech-to-speech models, where a large language model (LLM) acts as a backend providing candidate responses while the user is speaking. Instead of simulating the backend’s guidance, the approach derives guidance directly from the conversation corpus, using target responses for informative guidance and randomly sampled responses to simulate irrelevant updates. Experiments on synthetic dialogues and 3.8k hours of real conversations show that this technique yields response quality comparable to LLM-generated baselines while improving natural turn‑taking and audio‑judge naturalness.
By Manato Yaguchi, Yotaro Kubo, Hikaru Asano, So Kuroki
arXiv:2609.09852v1 Announce Type: new
Abstract: The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech intera...
By Heyang Liu, Jiayi Huang, Wenyang Xiao, Ziyang Cheng, Lixin Zhang, Zhen Liu, Miao He, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang
arXiv:2609.22697v1 Announce Type: new
Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated...
By Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue, Boyi Kang, Shilei Zhang, Jinglei Liu, Yue Wang, Junlan Feng, Bei Liu, Wei Xue
arXiv:2609.13076v1 Announce Type: cross
Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-en...
By Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath