arXiv Computation and Language

A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis

The paper introduces a native-reference phone‑class geometry that measures second‑language pronunciation deviation without needing pronunciation labels, read‑aloud prompts, or matched native recordings. By averaging self‑supervised representations for each phone‑class in a native speech corpus and applying singular value decomposition, the authors create a compact coordinate system. Projecting L2 utterances into this space, they find that distances to native coordinates correlate negatively with holistic speaking proficiency and pronunciation quality, indicating the geometry captures relevant acoustic‑phonetic information for spontaneous L2 speech.

arXiv AI
Sep 4

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

The paper presents a method for creating a compact fixed‑voice Thai text‑to‑speech system by training a student model on synthetic speech generated from a large voice‑cloning teacher. By using only a short 15‑second voice reference and carefully filtering synthetic data, the authors build an 82‑million‑parameter model, Wayu‑Paxa‑TTS‑Edge, that runs on device without reference audio. The system achieves strong performance—68.2 % challenge‑set keyword accuracy, 91.4 % pause precision, and low character error rates—while outperforming its teacher and approaching the quality of a larger Gemini 3.1 model.

By Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut
arXiv AI
Sep 12

Flexible and Interpretable Accent Distance Measurements

The paper introduces a method for measuring accent differences that balances interpretability and practicality. It proposes using articulatory representations obtained via articulatory inversion as an interpretable basis for accent comparison, while employing optimal transport to compare accents across any type of recording. This approach aims to overcome the limitations of traditional phonetic analyses and embedding‑based methods, which are either time‑consuming or non‑interpretable.

By Charles McGhee, Mark J. F. Gales, Kate M. Knill
arXiv Machine Learning
Jul 14

An Empirical Recipe for Universal Phone Recognition

arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.

By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv Machine Learning
Jun 30

BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

arXiv:2509. 15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences.

By Th\'eo Charlot, Tarek Kunze, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, Marvin Lavechin
arXiv AI
Aug 11

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

arXiv:2608. 09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.

By Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols