Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
arXiv:2604. 23354v3 Announce Type: replace-cross Abstract: Neural networks can be trained to learn task-relevant representations from data.
arXiv:2604. 23354v3 Announce Type: replace-cross Abstract: Neural networks can be trained to learn task-relevant representations from data.
The paper introduces the concept of second‑order patterns—latent structures that underlie a speaker recognition network’s classification of utterances. Using hierarchical clustering, the authors identify these patterns and interpret them with the HCCM method. They then propose a new task, second‑order pattern recognition, and present the HCNA method, which improves performance by matching unseen utterances to the extrapolation space of identified clusters.
arXiv:2609.39162v1 Announce Type: cross Abstract: Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermed...
arXiv:2607. 05276v1 Announce Type: cross Abstract: Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed speech segment to an x-vector, which is then used for downstream applications.
arXiv:2403.14830v2 Announce Type: replace Abstract: Deep clustering partitions complex high-dimensional data using deep neural networks for clustering. It involves projecting data into lower-dimensio...
arXiv:2405. 12775v2 Announce Type: replace-cross Abstract: Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.
arXiv:2606. 29335v1 Announce Type: cross Abstract: Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training and testing conditions.
arXiv:2606. 06740v1 Announce Type: cross Abstract: Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speaker speech generation.
arXiv:2609.09628v1 Announce Type: cross Abstract: Inferring speaker relationships from spoken conversations is an important step towards socially aware speech understanding. However, this task remain...
arXiv:2608. 06300v1 Announce Type: new Abstract: Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age.
arXiv:2511. 05913v2 Announce Type: replace-cross Abstract: New intent discovery (NID) seeks to recognize both new and known intents from unlabeled user utterances, which finds prevalent use in practical dialogue systems.
arXiv:2606. 17416v1 Announce Type: cross Abstract: Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become entangled with linguistic characteristics, degrading generalization across languages.