arXiv:2605. 20712v2 Announce Type: replace-cross Abstract: Automatic speech recognition replaces typing only when correction costs less than manual entry - a threshold determined by error types, not counts: fixing a misrecognized domain term costs far more than inserting a comma.
By Kavya Manohar, Arghya Bhattacharya, Kush Juvekar, Kumarmanas Nethil
SuTRA (Structurally-Unified Tokenization with Root Awareness) is a morphology-aware tokenization algorithm designed to address the problem of Morphological Shattering in morphologically rich Indic languages. It preserves the indivisibility of aksharas—complex orthographic syllables—by penalizing merges that cross morphological boundaries, thereby reducing over-fragmentation of words. The authors also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati, and demonstrate that SuTRA improves morphological alignment by up to 14.7% and semantic recoverability by 34% over BPE, leading to an average machine translation gain of +8.08 chrF2.
By Vaibhav Rathore, Siddhant Gole, Dadhichi Telwadkar, Rooshil Bhatia, Maulik Ruparel, Siddharth Surekha, Neha Bhargava
arXiv:2607. 23808v1 Announce Type: cross Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India.
By Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi, Kaushal Bhogale, Mitesh M. Khapra
UNESCO considers the Assyrian (Syriac) language an endangered language. Although Assyrians speak the language worldwide, the speaking population is uncertain (ranging from 500,000 to 1,500,000). Syria...
arXiv:2609.14967v1 Announce Type: new
Abstract: Quranic text is distributed in two orthographic forms that are byte-level distinct: the Uthmani script used in every printed mushaf, and the Standard (...
By Yahya Mohamed Elnawasany
The paper presents a multi‑stage framework for recognizing Kuzushiji characters in Japanese historical documents. It combines character detection, cropping, classification, reading‑order reconstruction via adaptive column clustering, and large‑language‑model‑based post‑OCR correction. The authors also augment data synthetically, correct dataset annotations, and introduce new test sets, achieving significant character error rate reductions on real, synthetic, and out‑of‑domain data.
By Rui-Yang Ju, Kohei Yamashita, Hirotaka Kameko, Shinsuke Mori
IndicDetect is a benchmark for evaluating AI‑generated text detection in Hindi, Telugu, and Tamil. It pairs curated human‑written texts with LLM‑generated counterparts across multiple domains and generators, testing detectors under domain shift, generator shift, and adversarial perturbation. The study shows that supervised neural detectors fail to generalize to unseen generators and attacks, with Hindi experiencing the greatest degradation, indicating that robustness—not peak accuracy—is the main weakness in Indic language detectors.
By Bhaskar Ganesh Devalla, Junchao Wu, Nilesh Dokuparthi, Greeshma Yaluru, Tatiana Muniz Rodriguez, Lidia S. Chao, Derek F. Wong
Large language models (LLMs) demonstrate strong Helpfulness, Harmlessness, and Honesty (3H) alignment in English-centric settings, but these gains transfer poorly to low-resource languages due to cult...
AraMS-28k is the largest publicly released line‑level dataset of genuine historical Arabic manuscripts, containing 14 books, 3,043 pages, and 28,600 annotated text lines (27,971 main‑text and 629 margin). The dataset spans three script traditions—Naskh, Ruq'ah, and Maghrebi—and includes a lithographed printed edition for format diversity. Each line is labeled as main‑text or margin, with margin lines that have a clear attachment point annotated with an insertion anchor to recover the manuscript’s true non‑linear reading order; both fully vocalized and diacritic‑normalized transcriptions are provided, and the data was produced via the RefLAM pipeline combining OCR, clean transcriptions, and human review.
"whyItMatters":"The dataset’s comprehensive line‑level annotations, including reading‑order anchors and dual transcription formats, enable reproducible research on Arabic manuscript recognition, layout analysis, and reading‑order recovery under a CC BY‑NC‑SA 4.0 license."
By Mohamed Guechaoui, Mohamed Diaa Zellagui, Souleyman Chaib, Sahraoui Dhelim
arXiv:2606. 25365v2 Announce Type: replace-cross Abstract: We present a study on low-resource machine translation for the Tangkhul-English (nmf-en) language pair.
By Chormi Zimik Vashai, Agniva Maiti
arXiv:2609.17539v1 Announce Type: new
Abstract: We present MudawanSn, a gold-standard resource of 1,271 sentence-aligned pairs manually translated from Wolof into Modern Standard Arabic (MSA). The so...
By Mouhamed Mbaye, Thierno Diop
The paper introduces MoirfEolas, a dataset of over 35,000 Irish words annotated with their morphological components, and CríochScore, a metric that measures how well tokenization aligns with these morphological boundaries. Using CríochScore, the authors evaluate common tokenization algorithms and find that the Unigram Language Model best aligns with Irish morphology. They also discuss trade‑offs between morphological alignment, compression, and vocabulary efficiency, offering practical guidance for Irish NLP development.
By Jane Adkins, Abigail Walsh, Brian Davis, Elaine U\'i Dhonnchadha