arXiv AI By MinJu Jeon, Younghan Park, Han Sung Park, Jong-Hwan Kim, Dong-Jin Kim, Hoyeon Lee

Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach

Read the original on arXiv AI →

The paper introduces UGTPhon, a grapheme-to-phoneme benchmark for user‑generated text in English, Vietnamese, and Korean, and presents a taxonomy for diagnosing pronunciation errors. It shows that existing G2P models and large language models struggle with canonical‑to‑non‑canonical text, with errors up to 66.8 PER points. A compositional G2P approach that uses exact‑match lookup and staged decoding reduces these errors and performs competitively with larger few‑shot LLMs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 18

Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data

The paper introduces a context‑aware neural grapheme‑to‑phoneme (G2P) system for unsegmented languages like Japanese, using a discriminative conditional random field over a word lattice built from dictionaries. It addresses data scarcity by generating over two million sentences with large language models. Experiments show the method surpasses traditional morphological analyzers and neural sequence models, achieving 99.62% target word reading accuracy and very low phoneme error rates on the Joyo‑Kanji‑Yomi benchmark.

By Rui Hu, Zhenpeng Zhan, Xiaolong Lin
arXiv Computation and Language
5d ago

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
Hugging Face Trending Papers
Sep 17

Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data

The paper introduces a context‑aware neural grapheme‑to‑phoneme (G2P) system for unsegmented languages like Japanese, combining a discriminative conditional random field over a dictionary‑based word lattice with large language model‑generated training data. By generating over two million synthetic sentences, the method addresses data scarcity and achieves superior performance compared to traditional morphological analyzers and neural sequence models. On the Joyo‑Kanji‑Yomi benchmark, it attains 99.62% target word reading accuracy, 0.32% target word phoneme error rate, and 0.14% sentence phoneme error rate.

arXiv Computation and Language
2d ago

Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring

The paper introduces a training‑free speech‑and‑text‑to‑pronunciation (ST2P) pipeline that combines lexical candidates from G2P tools with acoustic rescoring using frozen pretrained S2P models. By performing a left‑to‑right greedy search over whole‑sequence negative log‑likelihoods, the method achieves a dramatic reduction in character error rate on Japanese corpora, outperforming both baseline G2P/S2P approaches and commercial multimodal LLMs. The approach is also significantly faster—3–3.5× faster than beam search and twice as fast as direct decoding—while maintaining high accuracy across multiple languages.

By Hikaru Asano, Yotaro Kubo, So Kuroki