arXiv AI

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

arXiv:2606. 06740v1 Announce Type: cross Abstract: Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speaker speech generation.

arXiv Computation and Language
Sep 25

BiMamba2 Masked Discrete-Unit Prediction for Multilingual Speech Representation for Unsupervised Speech in the Wild Challenge

The paper presents BiMamba2, a 47.88‑million‑parameter bidirectional Mamba‑2 encoder trained with masked discrete‑unit prediction for multilingual speech representation. It was trained on 250 hours of unlabeled speech from 67 languages and evaluated in the Unsupervised Speech in the Wild Challenge, achieving an Adjusted Rand Index of 0.735 for speaker clustering while reporting lower performance on language identification and character error rate compared to supervised baselines. The authors also discuss a discrepancy between local‑official metric scales and checkpoint rankings, underscoring the limits of in‑distribution diagnostics for predicting Dynabench probe outcomes.

By Prakriti Subedi, Howard Prioleau, Saurav K Aryal
arXiv Computation and Language
Sep 25

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
arXiv Machine Learning
Jul 14

An Empirical Recipe for Universal Phone Recognition

arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.

By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv Computation and Language
Sep 18

Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models

The paper evaluates bias in phoneme-based automatic speech recognition (ASR) systems, focusing on WhisperIPA and ZIPA, which produce International Phonetic Alphabet (IPA) transcriptions. Using multilingual speech corpora and demographically annotated English datasets, the authors compare model-generated IPA against grapheme-to-phoneme (G2P) outputs with both standard phoneme error rate (PER) and a new Soft PER metric that allows linguistically similar substitutions. The study finds persistent disparities across language, gender, accent, ethnicity, and age, even when accounting for acceptable phonemic variation.

By Maneesha Rani Saha, Catherine Bao, Neal Patwari
arXiv Computation and Language
Sep 25

A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis

The paper introduces a native-reference phone‑class geometry that measures second‑language pronunciation deviation without needing pronunciation labels, read‑aloud prompts, or matched native recordings. By averaging self‑supervised representations for each phone‑class in a native speech corpus and applying singular value decomposition, the authors create a compact coordinate system. Projecting L2 utterances into this space, they find that distances to native coordinates correlate negatively with holistic speaking proficiency and pronunciation quality, indicating the geometry captures relevant acoustic‑phonetic information for spontaneous L2 speech.

By Tina Raissi, Nhan Phan, Chenxiao Wang, Mikko Kurimo
arXiv Machine Learning
Jun 30

BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

arXiv:2509. 15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences.

By Th\'eo Charlot, Tarek Kunze, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, Marvin Lavechin