The paper introduces UGTPhon, a grapheme-to-phoneme benchmark for user‑generated text in English, Vietnamese, and Korean, and presents a taxonomy for diagnosing pronunciation errors. It shows that existing G2P models and large language models struggle with canonical‑to‑non‑canonical text, with errors up to 66.8 PER points. A compositional G2P approach that uses exact‑match lookup and staged decoding reduces these errors and performs competitively with larger few‑shot LLMs.
By MinJu Jeon, Younghan Park, Han Sung Park, Jong-Hwan Kim, Dong-Jin Kim, Hoyeon Lee
arXiv:2606. 16019v1 Announce Type: cross Abstract: Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech.
By Alexander Metzger, Aruna Srivastava, Ruslan Mukhamedvaleev
The paper introduces a context‑aware neural grapheme‑to‑phoneme (G2P) system for unsegmented languages like Japanese, using a discriminative conditional random field over a word lattice built from dictionaries. It addresses data scarcity by generating over two million sentences with large language models. Experiments show the method surpasses traditional morphological analyzers and neural sequence models, achieving 99.62% target word reading accuracy and very low phoneme error rates on the Joyo‑Kanji‑Yomi benchmark.
By Rui Hu, Zhenpeng Zhan, Xiaolong Lin
The paper introduces a context‑aware neural grapheme‑to‑phoneme (G2P) system for unsegmented languages like Japanese, combining a discriminative conditional random field over a dictionary‑based word lattice with large language model‑generated training data. By generating over two million synthetic sentences, the method addresses data scarcity and achieves superior performance compared to traditional morphological analyzers and neural sequence models. On the Joyo‑Kanji‑Yomi benchmark, it attains 99.62% target word reading accuracy, 0.32% target word phoneme error rate, and 0.14% sentence phoneme error rate.
DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.
By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
The paper introduces a training‑free speech‑and‑text‑to‑pronunciation (ST2P) pipeline that combines lexical candidates from G2P tools with acoustic rescoring using frozen pretrained S2P models. By performing a left‑to‑right greedy search over whole‑sequence negative log‑likelihoods, the method achieves a dramatic reduction in character error rate on Japanese corpora, outperforming both baseline G2P/S2P approaches and commercial multimodal LLMs. The approach is also significantly faster—3–3.5× faster than beam search and twice as fast as direct decoding—while maintaining high accuracy across multiple languages.
By Hikaru Asano, Yotaro Kubo, So Kuroki
arXiv:2609.09889v1 Announce Type: new
Abstract: Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR m...
By Yonghyun Jun, Jimin Lee, Hwan Chang, Dongho Shin, Seolah Kim, Hwanhee Lee
arXiv:2609.09974v1 Announce Type: new
Abstract: Grapheme-to-phoneme conversion (G2P) refers to the task of converting a sequence of graphemes to a corresponding sequence of phonemes. While Filipino G...
By Lorenz Bernard Marqueses, Paulo Grane Gabriel Silva, Chastine Cabatay, Ericson Adler Tan, Ann Franchesca Laguna
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary.
The paper introduces a unified phoneme‑based TTS‑to‑ASR augmentation pipeline that uses a multilingual TTS model with language‑ID conditioning and incorporates grapheme‑to‑phoneme conversion, reference‑speech filtering, and candidate‑text selection. It proposes phoneme‑frequency‑guided selection (PFGS) to rank sentences based on phoneme frequencies from real ASR labels, and demonstrates that random augmentation and PFGS both improve ASR performance across Arabic, French, Italian, and Portuguese test sets, with PFGS yielding up to a 19.3% relative WER reduction. The study also shows that filtering reference speech can further lower WER by up to 0.59 points on certain datasets.
By Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang
arXiv:2510.22172v2 Announce Type: replace-cross
Abstract: The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanis...
By Ruixiang Mao, Xiangnan Ma, Qing Yang, Ziming Zhu, Yucheng Qiao, Yuan Ge, Tong Xiao, Shengxiang Gao, Zhengtao Yu, Jingbo Zhu
arXiv:2607. 21332v1 Announce Type: cross Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties.
By Zhiheng Qian, Aini Li, Hai Hu, Liang Zhao