arXiv AI

Automated Pronunciation Evaluation for Korean Toddler Speech using Speech Diarization and Self-Supervised Learning

arXiv:2606. 10213v1 Announce Type: cross Abstract: Speech sound disorders affect approximately 44% of Korean pediatric communication disorder cases, yet automated assessment tools for Korean toddler speech remain underdeveloped.

arXiv Machine Learning
Jun 30

BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

arXiv:2509. 15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences.

By Th\'eo Charlot, Tarek Kunze, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, Marvin Lavechin
arXiv AI
Aug 25

Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

The paper proposes a linguistically structured multi‑task learning framework for recognizing non‑canonical phonemes by decomposing phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. A hierarchical architecture with task‑specific heads and a cross‑attention fusion module is combined with semi‑supervised Momentum Pseudo‑Labeling and a cascaded training strategy that gradually introduces articulatory tasks. Experiments on the L2‑ARCTIC dataset demonstrate significant improvements over baseline models and produce interpretable error patterns aligned with phonological feature structure.

By Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang, Anamika Ragu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
arXiv Machine Learning
Sep 18

Enabling automatic transcription of child-centered audio recordings from real-world environments

The paper presents a method to automatically identify utterances in child-centered daylong audio recordings that can be reliably transcribed by modern ASR systems, enabling accurate transcription of a substantial portion of the speech. On four English corpora, the approach achieves a median WER of 0% and a mean WER of 16% when transcribing 30% of the total speech, compared to a median WER of 52% when transcribing all speech. Word frequency distributions from the automatic transcripts correlate strongly with manual annotations (r = 0.94 overall, r = 0.99 for frequent words).

By Daniil Kocharov, Azarias Galama, Okko R\"as\"anen
arXiv Computation and Language
Sep 25

A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis

The paper introduces a native-reference phone‑class geometry that measures second‑language pronunciation deviation without needing pronunciation labels, read‑aloud prompts, or matched native recordings. By averaging self‑supervised representations for each phone‑class in a native speech corpus and applying singular value decomposition, the authors create a compact coordinate system. Projecting L2 utterances into this space, they find that distances to native coordinates correlate negatively with holistic speaking proficiency and pronunciation quality, indicating the geometry captures relevant acoustic‑phonetic information for spontaneous L2 speech.

By Tina Raissi, Nhan Phan, Chenxiao Wang, Mikko Kurimo
arXiv Computation and Language
Aug 24

Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing

The study demonstrates that self‑supervised speech embeddings can track how children’s speech converges toward adult patterns over development. Using HuBERT‑BASE embeddings extracted from over 925 hours of child‑caregiver recordings, the researchers found that the acoustic distance between a child’s vocalizations and those of their female caregiver decreased with the child’s hearing age, even after controlling for pitch and vocalization length. This single distance metric also correlated with several standardized speech and language assessments from infancy through preschoolhood, suggesting a scalable, language‑neutral way to monitor spoken language development in everyday settings.

By L. Choy, A. S. Khan, S. Patrizi, D. Ye, J. Gross, M. Cychosz
Hugging Face Trending Papers
Jun 23

Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant

Early identification of speech sound errors in children is often limited by access to specialists, motivating lightweight screening tools that can operate outside the clinic. We present a screening pipeline for Polish-speaking children focused on sibilant substitutions, coupling a wav2vec2-based CTC token recognizer with alignment-based error typing and a template-grounded caregiver assistant for screening, not diagnosis.