arXiv Computation and Language By Seanghay Yath

THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer

Read the original on arXiv Computation and Language →

The paper introduces Tha, a Khmer text normalization and inverse text normalization toolkit that uses weighted finite-state transducers. Tha processes entire lines in a single shortest-path search and employs a second transducer to prevent token boundaries within Khmer syllables. On Google's Khmer test suite, Tha achieves perfect agreement on 274 cardinals and correctly rewrites 153 of 158 real TTS prompts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 1

Towards a Joint Khmer Text Recognition and Word Segmentation

The paper introduces a unified model that jointly performs Khmer text recognition and word segmentation, eliminating the need for a separate segmentation step. Using a connectionist-temporal-classification decoder, the model can be instructed to output Khmer text with or without word boundaries. Experiments across document, scene, and handwritten image datasets demonstrate that the model accurately recognizes characters and locates word boundaries, reducing error and latency compared to traditional sequential pipelines.

By Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing, Masakazu Iwamura, Koichi Kise
arXiv Computation and Language
Sep 18

A Phonemically Comprehensive, ASCII-Only Romanization Scheme for Thai and Lao: Systematic Cross-Lingual Correspondence and Chinese-User-Friendly Design

The paper introduces an ASCII‑only romanization system that covers all phonemic distinctions in Thai and Lao, including segmental contrasts, vowel length, and lexical tone. It ensures one‑symbol‑one‑phoneme transparency and systematic cross‑lingual correspondence between the two languages, while also aligning with Pinyin and Jyutping where possible. The design prioritizes synchronic phonetic clarity, offers optional historical tone annotations, and results in a readable, keyboard‑friendly, machine‑processable representation useful for language learning and cross‑lingual speech processing.

By Zijie Zhang, Tan Lee
arXiv Computation and Language
Sep 4

Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition

Dual-Form ASR (DF-ASR) is a framework that unifies spoken-form ASR and semantics-aware written-form inverse text normalization (ITN) for Chinese speech recognition. It uses paired spoken- and written-form supervision generated and judged by a large language model, and introduces an ITN-MWER objective to penalize errors on normalization-sensitive spans. DF-ASR also employs a REQUIRE-ITN/FORBID-ITN protocol to separately evaluate required normalization and forbidden-span preservation, achieving superior performance over open-source ASR-ITN systems while maintaining prompt-level control between transcript forms.

By Fengrun Zhang, Li Fu, Wangjin Zhou, Lu Fan, Youzheng Wu, Xiaodong He