arXiv Machine Learning By Nikhil Navas, Sergio Chevtchenko, Talisson Damiao, Saeed Afshar

BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization

Read the original on arXiv Machine Learning →

BranchShine-CR is a 25‑million‑parameter model that transcribes multilingual speech into the International Phonetic Alphabet (IPA). It uses log‑mel features, a rotary‑position E‑Branchformer encoder, intermediate self‑conditioned CTC, and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances it achieves a 4.47 % IPA character error rate, a 22.3 % relative improvement over ZIPA‑CTC‑NS and outperforms a similarly sized NeMo Conformer baseline across all 41 language labels.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 14

An Empirical Recipe for Universal Phone Recognition

arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.

By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv Computation and Language
Sep 18

Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models

The paper evaluates bias in phoneme-based automatic speech recognition (ASR) systems, focusing on WhisperIPA and ZIPA, which produce International Phonetic Alphabet (IPA) transcriptions. Using multilingual speech corpora and demographically annotated English datasets, the authors compare model-generated IPA against grapheme-to-phoneme (G2P) outputs with both standard phoneme error rate (PER) and a new Soft PER metric that allows linguistically similar substitutions. The study finds persistent disparities across language, gender, accent, ethnicity, and age, even when accounting for acceptable phonemic variation.

By Maneesha Rani Saha, Catherine Bao, Neal Patwari
arXiv Computation and Language
Aug 28

Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study

The paper introduces a unified phoneme‑based TTS‑to‑ASR augmentation pipeline that uses a multilingual TTS model with language‑ID conditioning and incorporates grapheme‑to‑phoneme conversion, reference‑speech filtering, and candidate‑text selection. It proposes phoneme‑frequency‑guided selection (PFGS) to rank sentences based on phoneme frequencies from real ASR labels, and demonstrates that random augmentation and PFGS both improve ASR performance across Arabic, French, Italian, and Portuguese test sets, with PFGS yielding up to a 19.3% relative WER reduction. The study also shows that filtering reference speech can further lower WER by up to 0.59 points on certain datasets.

By Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang
arXiv Computation and Language
5d ago

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

The paper introduces a token‑level extension of Omni‑Temporal Classification (OTC) for automatic speech recognition, allowing unsupported tokens to be bypassed while preserving supervision for the rest of the word. Across 19 languages and three corpora, this token‑level OTC consistently outperforms standard CTC, achieving the lowest mean word error rate on every dataset and a 9.45% average relative WER reduction. A predictive‑entropy‑indexed schedule replaces epoch‑based relaxation, reducing training‑length dependence while maintaining performance.

By Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh