VoiceTrace introduces a new benchmark, VoiceTrace-Bench, for hybrid speech retrieval that combines a textual query specifying "what" to retrieve with a reference speech specifying "who" to retrieve. The authors propose a two‑stage framework: VoiceTrace‑Emb, which learns unified audio‑text embeddings for efficient large‑scale retrieval, and VoiceTrace‑Reranker, which fine‑grains relevance by jointly examining query‑candidate pairs. Experiments show VoiceTrace outperforms existing methods on both traditional semantic speech retrieval benchmarks and the new hybrid setting.
By Aaron Yee, Fengjie Lu, Jiarui Hai, Chenang Jiang, Helin Wang, Siwei Tu, Weitao You, Lingyun Sun
arXiv:2608.22872v2 Announce Type: new
Abstract: Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline...
By Zhenghua Bao
arXiv:2608.22872v1 Announce Type: new
Abstract: Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline...
By Zhenghua Bao
arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.
By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.
arXiv:2608. 19936v1 Announce Type: cross Abstract: Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities.
By Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr C{\l}apa, Jens Madsen, Panagiotis Tzirakis
The paper introduces a cause-aware error recovery framework for cascaded Automatic Speech Recognition – Large Language Model (ASR‑LLM) pipelines in Spoken Dialogue Systems. It replaces simple ASR confidence filtering with precision‑focused detectors that use deep ASR latent representations to classify token‑level errors into perception, comprehension, and deletion failures. This fine‑grained diagnosis enables the LLM to execute targeted, multi‑turn clarification strategies, leading to a more than two‑fold increase in recall on domain‑shift errors and significant reductions in word error rate and downstream task errors across varied accents, distortions, and domains.
By Yizhou Peng, Ziyang Ma, Changsong Liu, Yi-Wen Chao, Xie Chen, Eng Siong Chng
arXiv:2609.22214v1 Announce Type: new
Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...
By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
arXiv:2507.18061v4 Announce Type: replace-cross
Abstract: Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks pri...
By Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang, Jing Zhou, Hang Lv, Mengjie Du, Yaodong Song, Jie Lian, Jian Kang, Jie Li, Yongxiang Li
arXiv:2607. 05365v1 Announce Type: cross Abstract: Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech.
By Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar, Ashish Hallur, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez
The paper introduces a black-box membership inference attack framework tailored for fine-tuned text-to-speech models, addressing challenges in query generation and representation engineering. It evaluates five query types, finding recitation queries most effective, and uses multi-level speech embeddings with temporal alignment for fine-grained comparison. Experiments on CosyVoice2, F5-TTS, and XTTS-v2 trained on VCTK and British Dialect datasets show high privacy leakage, with speaker-level AUC above 0.80 and record-level AUC between 0.80 and 0.90.
By Kunlin Cai, Kaiyuan Zhang, Zihang Xiang, Jinghuai Zhang, Abeer Alwan, Fnu Suya, Yuan Tian