A Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in...
The paper introduces a native-reference phone‑class geometry that measures second‑language pronunciation deviation without needing pronunciation labels, read‑aloud prompts, or matched native recordings. By averaging self‑supervised representations for each phone‑class in a native speech corpus and applying singular value decomposition, the authors create a compact coordinate system. Projecting L2 utterances into this space, they find that distances to native coordinates correlate negatively with holistic speaking proficiency and pronunciation quality, indicating the geometry captures relevant acoustic‑phonetic information for spontaneous L2 speech.
arXiv:2606. 17835v1 Announce Type: cross Abstract: This study examines the extent to which the wav2vec2.
arXiv:2606. 11542v1 Announce Type: cross Abstract: Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations.
arXiv:2607. 09020v1 Announce Type: cross Abstract: Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately.
arXiv:2509. 15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences.