The paper introduces a unified model that jointly performs Khmer text recognition and word segmentation, eliminating the need for a separate segmentation step. Using a connectionist-temporal-classification decoder, the model can be instructed to output Khmer text with or without word boundaries. Experiments across document, scene, and handwritten image datasets demonstrate that the model accurately recognizes characters and locates word boundaries, reducing error and latency compared to traditional sequential pipelines.
By Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing, Masakazu Iwamura, Koichi Kise
arXiv:2609.21084v1 Announce Type: cross
Abstract: Modern automatic speech recognition (ASR) systems trained on extremely large datasets can produce transcripts with numbers written in Arabic numerals...
By Stanis{\l}aw Kacprzak, Mieszko Fra\'s
arXiv:2609.14542v1 Announce Type: new
Abstract: Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utt...
By Ahmad Amirivojdan, Farzad Nadiri, Abolfazl Alizadeh, Shaghayegh Yaraghi
arXiv:2608.28635v1 Announce Type: cross
Abstract: Recent multimodal large language models (MLLMs) have advanced document understanding, visual question answering, and text extraction. However, their...
By Nimol Thuon, Panhapin Theang
The paper introduces an ASCII‑only romanization system that covers all phonemic distinctions in Thai and Lao, including segmental contrasts, vowel length, and lexical tone. It ensures one‑symbol‑one‑phoneme transparency and systematic cross‑lingual correspondence between the two languages, while also aligning with Pinyin and Jyutping where possible. The design prioritizes synchronic phonetic clarity, offers optional historical tone annotations, and results in a readable, keyboard‑friendly, machine‑processable representation useful for language learning and cross‑lingual speech processing.
By Zijie Zhang, Tan Lee
Dual-Form ASR (DF-ASR) is a framework that unifies spoken-form ASR and semantics-aware written-form inverse text normalization (ITN) for Chinese speech recognition. It uses paired spoken- and written-form supervision generated and judged by a large language model, and introduces an ITN-MWER objective to penalize errors on normalization-sensitive spans. DF-ASR also employs a REQUIRE-ITN/FORBID-ITN protocol to separately evaluate required normalization and forbidden-span preservation, achieving superior performance over open-source ASR-ITN systems while maintaining prompt-level control between transcript forms.
By Fengrun Zhang, Li Fu, Wangjin Zhou, Lu Fan, Youzheng Wu, Xiaodong He
arXiv:2510.10774v4 Announce Type: replace-cross
Abstract: Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech...
By Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery
arXiv:2608. 12327v1 Announce Type: cross Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol.
By Suman Paudel, Sarbin Sayami
BanglaKontho is a newly released 20‑hour single‑speaker Bangla text‑to‑speech corpus derived from professional audiobook recordings, comprising 7,050 segmented utterances with verified transcripts at 24 kHz. The project also provides a reusable Bangla text normalizer that handles Bangladeshi‑style digit grouping, currency and date expressions, Danda punctuation, and Unicode normalization, along with the full preprocessing pipeline. A baseline MB‑iSTFT‑VITS model trained from scratch on this corpus achieves a 9.5 % WER and 4.46 naturalness MOS, outperforming the same architecture retrained on the smaller 12‑hour IndicTTS‑Bn corpus.
By Mizbaul Haque Maruf
arXiv:2609.24310v1 Announce Type: cross
Abstract: Text-to-speech models for Bantu tonal languages are challenged by a tonal system that is rooted in both the lexis (i.e., the inventory of words, stem...
By Antoine Nzeyimana
The paper shows that HuggingFace’s ByteLevel pre‑tokenizer, which treats a word as a sequence of Unicode letters, splits abugida scripts at every vowel sign, creating a training‑free lower bound on tokenizer fertility. Across 26 languages, all 17 abugidas exhibit increased token counts (up to 9×), while Latin, Cyrillic, Hangul, and Han remain unchanged. The authors demonstrate that correcting the character class reduces Nepali token counts, improves model performance, and that this issue is widespread in popular HuggingFace models.
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
arXiv:2605. 20712v2 Announce Type: replace-cross Abstract: Automatic speech recognition replaces typing only when correction costs less than manual entry - a threshold determined by error types, not counts: fixing a misrecognized domain term costs far more than inserting a comma.
By Kavya Manohar, Arghya Bhattacharya, Kush Juvekar, Kumarmanas Nethil