arXiv AI By Naman Kothari, Arjun Gangwar, Adarsh Arigala, S Umesh

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

Read the original on arXiv AI →

arXiv:2606. 06740v1 Announce Type: cross Abstract: Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speaker speech generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 25

BiMamba2 Masked Discrete-Unit Prediction for Multilingual Speech Representation for Unsupervised Speech in the Wild Challenge

The paper presents BiMamba2, a 47.88‑million‑parameter bidirectional Mamba‑2 encoder trained with masked discrete‑unit prediction for multilingual speech representation. It was trained on 250 hours of unlabeled speech from 67 languages and evaluated in the Unsupervised Speech in the Wild Challenge, achieving an Adjusted Rand Index of 0.735 for speaker clustering while reporting lower performance on language identification and character error rate compared to supervised baselines. The authors also discuss a discrepancy between local‑official metric scales and checkpoint rankings, underscoring the limits of in‑distribution diagnostics for predicting Dynabench probe outcomes.

By Prakriti Subedi, Howard Prioleau, Saurav K Aryal
arXiv Computation and Language
Sep 25

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
arXiv Machine Learning
Jul 14

An Empirical Recipe for Universal Phone Recognition

arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.

By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen