arXiv Computation and Language By L. Choy, A. S. Khan, S. Patrizi, D. Ye, J. Gross, M. Cychosz

Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing

Read the original on arXiv Computation and Language →

The study demonstrates that self‑supervised speech embeddings can track how children’s speech converges toward adult patterns over development. Using HuBERT‑BASE embeddings extracted from over 925 hours of child‑caregiver recordings, the researchers found that the acoustic distance between a child’s vocalizations and those of their female caregiver decreased with the child’s hearing age, even after controlling for pitch and vocalization length. This single distance metric also correlated with several standardized speech and language assessments from infancy through preschoolhood, suggesting a scalable, language‑neutral way to monitor spoken language development in everyday settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jun 30

BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

arXiv:2509. 15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences.

By Th\'eo Charlot, Tarek Kunze, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, Marvin Lavechin
arXiv Machine Learning
Sep 18

Enabling automatic transcription of child-centered audio recordings from real-world environments

The paper presents a method to automatically identify utterances in child-centered daylong audio recordings that can be reliably transcribed by modern ASR systems, enabling accurate transcription of a substantial portion of the speech. On four English corpora, the approach achieves a median WER of 0% and a mean WER of 16% when transcribing 30% of the total speech, compared to a median WER of 52% when transcribing all speech. Word frequency distributions from the automatic transcripts correlate strongly with manual annotations (r = 0.94 overall, r = 0.99 for frequent words).

By Daniil Kocharov, Azarias Galama, Okko R\"as\"anen
arXiv Machine Learning
Aug 13

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

arXiv:2608. 11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts.

By Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan, Kexin Hu, Bashima Islam, Mark Hasegawa-Johnson, Nancy L. McElwain
Hugging Face Trending Papers
Aug 12

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers.