arXiv:2609.15203v1 Announce Type: cross
Abstract: Speaker recognition neural networks learn latent representations (i.e. speaker embeddings) from input utterances to recognise speaker identities. How...
By Yanze Xu, Wenwu Wang, Mark D. Plumbley
The paper introduces the concept of second‑order patterns—latent structures that underlie a speaker recognition network’s classification of utterances. Using hierarchical clustering, the authors identify these patterns and interpret them with the HCCM method. They then propose a new task, second‑order pattern recognition, and present the HCNA method, which improves performance by matching unseen utterances to the extrapolation space of identified clusters.
By Yanze Xu, Wenwu Wang, Mark D. Plumbley
arXiv:2609.39162v1 Announce Type: cross
Abstract: Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermed...
By Yehoshua Dissen, Joseph Keshet, Eduard Golshtein
arXiv:2606. 06740v1 Announce Type: cross Abstract: Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speaker speech generation.
By Naman Kothari, Arjun Gangwar, Adarsh Arigala, S Umesh
arXiv:2405. 12775v2 Announce Type: replace-cross Abstract: Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.
By Hanlei Zhang, Hua Xu, Fei Long, Xin Wang, Kai Gao
arXiv:2609.23194v1 Announce Type: new
Abstract: Acoustic representation learning is crucial for speech processing, yet low-resource languages (LRLs) face severe data scarcity, limiting the effectiven...
By Yannick Yomie Nzeuhang, Paulin Melatagia Yonta, Marie Tahon
arXiv:2608. 06990v1 Announce Type: cross Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning.
By Yuning Yu, Jos\'e Rodr\'iguez-Pi\~neiro, Xuefeng Yin, Bin Feng
arXiv:2403.14830v2 Announce Type: replace
Abstract: Deep clustering partitions complex high-dimensional data using deep neural networks for clustering. It involves projecting data into lower-dimensio...
By Zeya Wang, Chenglong Ye
arXiv:2608. 06300v1 Announce Type: new Abstract: Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age.
By Arya Labroo, Mengjie Qian, Kate Knill
arXiv:2606. 29335v1 Announce Type: cross Abstract: Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training and testing conditions.
By Chuxiao Zuo, Yao Zhu, Minqiang Xu, Manhong Wang, Yunke Zhang, Fei Huang
The paper introduces the Echo Chamber Effect, a failure mode in Graph Neural Networks where intra-community representations collapse while inter-community separation remains, differing from traditional oversmoothing. It proposes the Echo Chamber Index (ECI) to detect this effect by stratifying pairwise distances by community membership. Building on this analysis, the authors present Community-Aware Split Propagation (CASP), a lightweight plugin that decouples intra- and inter-community aggregation and learns their balance from label structure, improving performance across various GNN backbones in both homophilic and heterophilic settings.
By Asela Hevapathige, Ahad N. Zehmakan, Asiri Wijesinghe, Saman Halgamuge
arXiv:2609.23191v1 Announce Type: new
Abstract: Speech-text space alignment is a multimodal representation learning method consisting to map different speech and text into a shared representation spa...
By Yannick Yomie Nzeuhang, Marie Tahon, Paulin Melatagia Yonta