arXiv:2609.15203v1 Announce Type: cross
Abstract: Speaker recognition neural networks learn latent representations (i.e. speaker embeddings) from input utterances to recognise speaker identities. How...
By Yanze Xu, Wenwu Wang, Mark D. Plumbley
The paper introduces the concept of second‑order patterns—latent structures that underlie a speaker recognition network’s classification of utterances. Using hierarchical clustering, the authors identify these patterns and interpret them with the HCCM method. They then propose a new task, second‑order pattern recognition, and present the HCNA method, which improves performance by matching unseen utterances to the extrapolation space of identified clusters.
By Yanze Xu, Wenwu Wang, Mark D. Plumbley
arXiv:2609.39162v1 Announce Type: cross
Abstract: Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermed...
By Yehoshua Dissen, Joseph Keshet, Eduard Golshtein
arXiv:2606. 06740v1 Announce Type: cross Abstract: Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speaker speech generation.
By Naman Kothari, Arjun Gangwar, Adarsh Arigala, S Umesh
arXiv:2405. 12775v2 Announce Type: replace-cross Abstract: Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.
By Hanlei Zhang, Hua Xu, Fei Long, Xin Wang, Kai Gao
arXiv:2609.23194v1 Announce Type: new
Abstract: Acoustic representation learning is crucial for speech processing, yet low-resource languages (LRLs) face severe data scarcity, limiting the effectiven...
By Yannick Yomie Nzeuhang, Paulin Melatagia Yonta, Marie Tahon