arXiv AI

Interpreting hierarchical organisation of speaker embeddings

arXiv AI
Sep 12

Exploring Second-Order Pattern Recognition in Speaker Recognition

The paper introduces the concept of second‑order patterns—latent structures that underlie a speaker recognition network’s classification of utterances. Using hierarchical clustering, the authors identify these patterns and interpret them with the HCCM method. They then propose a new task, second‑order pattern recognition, and present the HCNA method, which improves performance by matching unseen utterances to the extrapolation space of identified clusters.

By Yanze Xu, Wenwu Wang, Mark D. Plumbley
arXiv AI
Jul 7

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

arXiv:2607. 05276v1 Announce Type: cross Abstract: Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed speech segment to an x-vector, which is then used for downstream applications.

By Thomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez, Jesus Villalba Lopez, Najim Dehak
arXiv AI
Aug 7

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

arXiv:2608. 06300v1 Announce Type: new Abstract: Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age.

By Arya Labroo, Mengjie Qian, Kate Knill