arXiv Computation and Language
The paper introduces grounded glossary generation, a structured NLP task that asks models to recover semantically meaningful Sanskrit phrases and provide translation‑grounded meanings from a sloka‑translation pair, mirroring the traditional patha commentary practice. A benchmark of 31,316 sloka‑translation‑glossary triples from the Valmiki Ramayana and Srimad Bhagavatam is built, evaluated with Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Experiments with Gemma‑3n‑E4B, Gemma‑3‑12B, Phi‑4, and Qwen3.5‑9B show that instruction fine‑tuning outperforms prompting, and explicit segmentation further improves results, though over‑segmentation of sandhi and samasa compounds remains the main error source, highlighting morphological modeling as a key bottleneck.
arXiv:2608.23120v1 Announce Type: cross
Abstract: Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural l...
By Edawanbiang Dhar Surmila Thokchom, Thoudam Doren Singh
arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv:2606. 08272v1 Announce Type: cross Abstract: AgriGov is a curated, trilingual (English-Hindi-Marathi) dataset designed to address the scarcity of domain-grounded multilingual resources for agricultural policies and farmer welfare schemes.
By Mohsina Bilal, Gopakumar G
arXiv:2606. 24825v1 Announce Type: cross Abstract: Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing.
By Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Raviraj Joshi
arXiv:2606. 09767v1 Announce Type: cross Abstract: Neural machine translation for digitally low-resource Indigenous languages is often hindered by extreme data scarcity, prompting reliance on extractive web-scraping.
By Alexander Chulzhanov, Soeren Eberhardt, Arjun Mukherjee
arXiv:2402. 18121v2 Announce Type: replace-cross Abstract: This study assesses four cutting-edge language models in the underexplored Aminoacian language.
By Yunze Xiao, Yiyang Pan
arXiv:2606. 12392v1 Announce Type: cross Abstract: Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry.
By Haotao Xie
Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable. We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary.
Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry. However, domain-specific research on precise translation and affective-semantic understanding of classical poetry remains limited.
SuTRA (Structurally-Unified Tokenization with Root Awareness) is a morphology-aware tokenization algorithm designed to address the problem of Morphological Shattering in morphologically rich Indic languages. It preserves the indivisibility of aksharas—complex orthographic syllables—by penalizing merges that cross morphological boundaries, thereby reducing over-fragmentation of words. The authors also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati, and demonstrate that SuTRA improves morphological alignment by up to 14.7% and semantic recoverability by 34% over BPE, leading to an average machine translation gain of +8.08 chrF2.
By Vaibhav Rathore, Siddhant Gole, Dadhichi Telwadkar, Rooshil Bhatia, Maulik Ruparel, Siddharth Surekha, Neha Bhargava
arXiv:2011.03783v3 Announce Type: replace-cross
Abstract: In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with an...
By Lifeng Han, Najet Hadj Mohamed, Malak Rassem, Gareth Jones, Alan Smeaton, Goran Nenadic
arXiv:2605. 20712v2 Announce Type: replace-cross Abstract: Automatic speech recognition replaces typing only when correction costs less than manual entry - a threshold determined by error types, not counts: fixing a misrecognized domain term costs far more than inserting a comma.
By Kavya Manohar, Arghya Bhattacharya, Kush Juvekar, Kumarmanas Nethil