arXiv:2608.31037v1 Announce Type: new
Abstract: Neural audio codecs (NACs) convert speech into discrete token sequences, and prior work has reported that these sequences follow language-like statisti...
By Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando, Yuki Saito, Hiroshi Saruwatari
arXiv:2609.26306v1 Announce Type: new
Abstract: This work presents the task and results of the CHiME-9 challenge for Enhancing Conversations to address Hearing Impairment. The challenge considers the...
By Robert Sutherland, Thomas Kuebert, Marko Lugger, Stefan Petrausch, Eline Borch Petersen, Juan Azcarreta Ortiz, Buye Xu, Stefan Goetze, Jon Barker
arXiv:2606. 19951v1 Announce Type: cross Abstract: Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear.
By Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata
arXiv:2608. 02235v1 Announce Type: cross Abstract: Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages.
By Ali Jafar, Amal Sarmad, Shifa Yousaf, Maryam Bashir
arXiv:2608. 05727v1 Announce Type: cross Abstract: Neural Audio Codecs are widely adopted in speech generation and editing.
By June Young Yi, Dongwook Lee, Jiheum Yeom, Sungroh Yoon
arXiv:2608.21176v1 Announce Type: cross
Abstract: Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential for evalua...
By Naiyuan Li, Li Dong, Diqun Yan