arXiv Computation and Language By Zijie Zhang, Tan Lee

A Phonemically Comprehensive, ASCII-Only Romanization Scheme for Thai and Lao: Systematic Cross-Lingual Correspondence and Chinese-User-Friendly Design

Read the original on arXiv Computation and Language →

The paper introduces an ASCII‑only romanization system that covers all phonemic distinctions in Thai and Lao, including segmental contrasts, vowel length, and lexical tone. It ensures one‑symbol‑one‑phoneme transparency and systematic cross‑lingual correspondence between the two languages, while also aligning with Pinyin and Jyutping where possible. The design prioritizes synchronic phonetic clarity, offers optional historical tone annotations, and results in a readable, keyboard‑friendly, machine‑processable representation useful for language learning and cross‑lingual speech processing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
6d ago

THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer

The paper introduces Tha, a Khmer text normalization and inverse text normalization toolkit that uses weighted finite-state transducers. Tha processes entire lines in a single shortest-path search and employs a second transducer to prevent token boundaries within Khmer syllables. On Google's Khmer test suite, Tha achieves perfect agreement on 274 cardinals and correctly rewrites 153 of 158 real TTS prompts.

By Seanghay Yath
arXiv Computation and Language
Sep 18

Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models

The paper evaluates bias in phoneme-based automatic speech recognition (ASR) systems, focusing on WhisperIPA and ZIPA, which produce International Phonetic Alphabet (IPA) transcriptions. Using multilingual speech corpora and demographically annotated English datasets, the authors compare model-generated IPA against grapheme-to-phoneme (G2P) outputs with both standard phoneme error rate (PER) and a new Soft PER metric that allows linguistically similar substitutions. The study finds persistent disparities across language, gender, accent, ethnicity, and age, even when accounting for acceptable phonemic variation.

By Maneesha Rani Saha, Catherine Bao, Neal Patwari
arXiv Machine Learning
Jul 14

An Empirical Recipe for Universal Phone Recognition

arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.

By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv Computation and Language
Sep 21

Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

arXiv:2609. 21362v1 Announce Type: new Abstract: Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies.

By Nghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le, Phu Gia Hoang, Dat Tien Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
arXiv Computation and Language
Sep 7

Evaluation of Phonetic Encoding Algorithms on Transcription Datasets

The paper introduces a new evaluation framework for phonetic encoding algorithms, using a generalized Rand Index called the Hüllermeier‑Rifqi Index. It measures discordance by comparing pairwise similarity scores of ground‑truth IPA transcriptions with those of encoded strings, adjusted against a random generator. The method is applied to multilingual datasets, assessing recall via collision rate and demonstrating its use in evaluating orthographic transparency.

By Can \"Ozbey, Emre Kaplan, Berkin Deniz Kahya