arXiv:2604. 23354v3 Announce Type: replace-cross Abstract: Neural networks can be trained to learn task-relevant representations from data.
By Yanze Xu, Wenwu Wang, Mark D. Plumbley
arXiv:2609.15203v1 Announce Type: cross
Abstract: Speaker recognition neural networks learn latent representations (i.e. speaker embeddings) from input utterances to recognise speaker identities. How...
By Yanze Xu, Wenwu Wang, Mark D. Plumbley
arXiv:2609.23194v1 Announce Type: new
Abstract: Acoustic representation learning is crucial for speech processing, yet low-resource languages (LRLs) face severe data scarcity, limiting the effectiven...
By Yannick Yomie Nzeuhang, Paulin Melatagia Yonta, Marie Tahon
arXiv:2609.39162v1 Announce Type: cross
Abstract: Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermed...
By Yehoshua Dissen, Joseph Keshet, Eduard Golshtein
arXiv:2405. 12775v2 Announce Type: replace-cross Abstract: Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.
By Hanlei Zhang, Hua Xu, Fei Long, Xin Wang, Kai Gao
arXiv:2609.38887v1 Announce Type: cross
Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effect...
By Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna