arXiv:2608. 09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.
By Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols
arXiv:2608.21176v1 Announce Type: cross
Abstract: Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential for evalua...
By Naiyuan Li, Li Dong, Diqun Yan
arXiv:2605.28227v2 Announce Type: replace
Abstract: Speech translation models are increasingly capable of preserving speech-specific information (e.g., speaker gender, prosody, and emphasis), yet eva...
By Maike Z\"ufle, Danni Liu, Vil\'em Zouhar, Jan Niehues
arXiv:2606. 15888v1 Announce Type: cross Abstract: Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent.
By Jialong Mai, Jinxin Ji, Xiaofen Xing, Wencui Liu, Xiangmin Xu
arXiv:2606. 20137v1 Announce Type: cross Abstract: Existing mean opinion score (MOS) prediction models typically predict utterance-level naturalness MOS and can be insensitive to localized pitch-accent errors.
By Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu
arXiv:2609.35952v1 Announce Type: cross
Abstract: We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real hum...
By Shen Yan, Duc Le, Irina-Elena Veliche
arXiv:2609.13150v1 Announce Type: cross
Abstract: Reference-free quality predictors such as UTMOS, DNSMOS and SCOREQ are the de facto automatic evaluators for text-to-speech (TTS) and are increasingl...
By Antonis Asonitis, Juan Pablo Zuluaga Gomez, Francesco Verdini, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet
arXiv:2608. 02235v1 Announce Type: cross Abstract: Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages.
By Ali Jafar, Amal Sarmad, Shifa Yousaf, Maryam Bashir
The paper investigates whether audio large language models (Audio LLMs) can detect when their own transcriptions are unreliable. It finds that the models are poor at self-assessment and that existing methods offer limited detection. By leveraging audio-encoder representations, the authors develop a lightweight predictor that accurately flags unreliable transcriptions and can prompt user clarification without altering the underlying model.
By Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh
arXiv:2606. 19597v1 Announce Type: cross Abstract: Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences.
By Junyi Fan, Donald S. Williamson
arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.
By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained paralinguistic distinctions underexplored.