arXiv:2606. 11542v1 Announce Type: cross Abstract: Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations.
By Chihiro Taguchi, \'Eric Le Ferrand, Hirosi Nakagawa, Hitomi Ono, Kanji Kato, Emily Prud'hommeaux, David Chiang
arXiv:2609.28060v1 Announce Type: new
Abstract: Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for...
By Tina Raissi, Nhan Phan, Mikko Kurimo
arXiv:2607. 01965v1 Announce Type: cross Abstract: Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms.
By Sneha Ray Barman, Neeraj Kumar Sharma, Shakuntala Mahanta
arXiv:2601.09050v2 Announce Type: replace
Abstract: Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech represe...
By Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu
Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this.
Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in...