arXiv Machine Learning

Enabling automatic transcription of child-centered audio recordings from real-world environments

The paper presents a method to automatically identify utterances in child-centered daylong audio recordings that can be reliably transcribed by modern ASR systems, enabling accurate transcription of a substantial portion of the speech. On four English corpora, the approach achieves a median WER of 0% and a mean WER of 16% when transcribing 30% of the total speech, compared to a median WER of 52% when transcribing all speech. Word frequency distributions from the automatic transcripts correlate strongly with manual annotations (r = 0.94 overall, r = 0.99 for frequent words).

arXiv Machine Learning
Jun 30

BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

arXiv:2509. 15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences.

By Th\'eo Charlot, Tarek Kunze, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, Marvin Lavechin
arXiv Computation and Language
Aug 24

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.

By Kawshik Kumar Paul, Md. Nafiul Alam Fuji
arXiv Computation and Language
Sep 11

Whisper-Based Speech Transcription from Videos Across Multiple Languages for Cross-Cultural Understanding

The paper introduces methods to improve speech recognition for multilingual video transcription using Whisper-based tools, targeting cross‑cultural understanding. It reports an average transcription error rate of 30% across seven languages, which can be lowered to 20% with modest fine‑tuning. The authors also release associated speech and metadata to aid community refinement of these techniques.

By Michael Picheny
arXiv Computation and Language
Sep 25

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

The paper introduces a token‑level extension of Omni‑Temporal Classification (OTC) for automatic speech recognition, allowing unsupported tokens to be bypassed while preserving supervision for the rest of the word. Across 19 languages and three corpora, this token‑level OTC consistently outperforms standard CTC, achieving the lowest mean word error rate on every dataset and a 9.45% average relative WER reduction. A predictive‑entropy‑indexed schedule replaces epoch‑based relaxation, reducing training‑length dependence while maintaining performance.

By Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh
arXiv Computation and Language
Sep 24

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

The paper introduces a trilingual spoken hallucination detection benchmark covering English, Russian, and Kazakh news, with 12,013 samples that include synthetic alterations and severity levels, as well as 290 fact‑checked misinformation items. Detectors are evaluated in a reference‑free setting on text, ASR transcripts, and audio, revealing that most models underperform baseline classifiers, except Gemma‑3n on transcripts. Synthetic‑trained detectors achieve high macro‑F1 scores on real‑world misinformation, but Russian provenance analysis highlights model‑dependent signals that confound synthetic benchmarks.

By Meruyert Aristombayeva, Jason S. Lucas, Chaewan Chun, Dongwon Lee