Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
arXiv:2604. 23354v3 Announce Type: replace-cross Abstract: Neural networks can be trained to learn task-relevant representations from data.
The paper introduces the concept of second‑order patterns—latent structures that underlie a speaker recognition network’s classification of utterances. Using hierarchical clustering, the authors identify these patterns and interpret them with the HCCM method. They then propose a new task, second‑order pattern recognition, and present the HCNA method, which improves performance by matching unseen utterances to the extrapolation space of identified clusters.
arXiv:2604. 23354v3 Announce Type: replace-cross Abstract: Neural networks can be trained to learn task-relevant representations from data.
arXiv:2609.15203v1 Announce Type: cross Abstract: Speaker recognition neural networks learn latent representations (i.e. speaker embeddings) from input utterances to recognise speaker identities. How...
arXiv:2609.23194v1 Announce Type: new Abstract: Acoustic representation learning is crucial for speech processing, yet low-resource languages (LRLs) face severe data scarcity, limiting the effectiven...
arXiv:2609.39162v1 Announce Type: cross Abstract: Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermed...
arXiv:2405. 12775v2 Announce Type: replace-cross Abstract: Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.
arXiv:2609.38887v1 Announce Type: cross Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effect...
arXiv:2606. 14647v1 Announce Type: cross Abstract: Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their predictions remain difficult to interpret.
arXiv:2607. 08111v1 Announce Type: cross Abstract: Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable.
arXiv:2602. 05670v2 Announce Type: replace-cross Abstract: Advances in AIGC technologies have enabled the synthesis of highly realistic audio deepfakes capable of deceiving human auditory perception.
arXiv:2603. 01006v3 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth.
arXiv:2607. 22646v1 Announce Type: new Abstract: Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several candidates without consensus, and none has been grounded in the model's internal activations.
arXiv:2605. 00025v3 Announce Type: replace-cross Abstract: Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communication for individuals with speech-impairing conditions.