arXiv Computation and Language

THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer

The paper introduces Tha, a Khmer text normalization and inverse text normalization toolkit that uses weighted finite-state transducers. Tha processes entire lines in a single shortest-path search and employs a second transducer to prevent token boundaries within Khmer syllables. On Google's Khmer test suite, Tha achieves perfect agreement on 274 cardinals and correctly rewrites 153 of 158 real TTS prompts.

arXiv Computation and Language
Sep 1

Towards a Joint Khmer Text Recognition and Word Segmentation

The paper introduces a unified model that jointly performs Khmer text recognition and word segmentation, eliminating the need for a separate segmentation step. Using a connectionist-temporal-classification decoder, the model can be instructed to output Khmer text with or without word boundaries. Experiments across document, scene, and handwritten image datasets demonstrate that the model accurately recognizes characters and locates word boundaries, reducing error and latency compared to traditional sequential pipelines.

By Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing, Masakazu Iwamura, Koichi Kise
arXiv Computation and Language
Sep 18

A Phonemically Comprehensive, ASCII-Only Romanization Scheme for Thai and Lao: Systematic Cross-Lingual Correspondence and Chinese-User-Friendly Design

The paper introduces an ASCII‑only romanization system that covers all phonemic distinctions in Thai and Lao, including segmental contrasts, vowel length, and lexical tone. It ensures one‑symbol‑one‑phoneme transparency and systematic cross‑lingual correspondence between the two languages, while also aligning with Pinyin and Jyutping where possible. The design prioritizes synchronic phonetic clarity, offers optional historical tone annotations, and results in a readable, keyboard‑friendly, machine‑processable representation useful for language learning and cross‑lingual speech processing.

By Zijie Zhang, Tan Lee
arXiv Computation and Language
Sep 4

Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition

Dual-Form ASR (DF-ASR) is a framework that unifies spoken-form ASR and semantics-aware written-form inverse text normalization (ITN) for Chinese speech recognition. It uses paired spoken- and written-form supervision generated and judged by a large language model, and introduces an ITN-MWER objective to penalize errors on normalization-sensitive spans. DF-ASR also employs a REQUIRE-ITN/FORBID-ITN protocol to separately evaluate required normalization and forbidden-span preservation, achieving superior performance over open-source ASR-ITN systems while maintaining prompt-level control between transcript forms.

By Fengrun Zhang, Li Fu, Wangjin Zhou, Lu Fan, Youzheng Wu, Xiaodong He
arXiv Computation and Language
Sep 25

BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech

BanglaKontho is a newly released 20‑hour single‑speaker Bangla text‑to‑speech corpus derived from professional audiobook recordings, comprising 7,050 segmented utterances with verified transcripts at 24 kHz. The project also provides a reusable Bangla text normalizer that handles Bangladeshi‑style digit grouping, currency and date expressions, Danda punctuation, and Unicode normalization, along with the full preprocessing pipeline. A baseline MB‑iSTFT‑VITS model trained from scratch on this corpus achieves a 9.5 % WER and 4.46 naturalness MOS, outperforming the same architecture retrained on the smaller 12‑hour IndicTTS‑Bn corpus.

By Mizbaul Haque Maruf
arXiv Machine Learning
Aug 28

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

The paper shows that HuggingFace’s ByteLevel pre‑tokenizer, which treats a word as a sequence of Unicode letters, splits abugida scripts at every vowel sign, creating a training‑free lower bound on tokenizer fertility. Across 26 languages, all 17 abugidas exhibit increased token counts (up to 9×), while Latin, Cyrillic, Hangul, and Han remain unchanged. The authors demonstrate that correcting the character class reduces Nepali token counts, improves model performance, and that this issue is widespread in popular HuggingFace models.

By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun