The study examines Urdu light verbs, which add schematic event meaning while staying lexically linked to their main verbs. Using contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT on 1,126 sentences, the authors find significant representational separation between main and light uses across all verb–model pairs, yet main and light uses of the same lemma remain closer than mismatched pairs. In a seven‑way prediction task limited to light uses, UrduBERT achieves 0.866 accuracy and 0.852 macro‑F1, and maintains 0.782 accuracy when tested on unseen preceding forms, demonstrating generalization beyond local verb combinations.
By Farah Adeeba, Miriam Butt
The paper introduces a new benchmark for Urdu‑to‑English idiomatic translation, featuring 4,000 manually verified sentence pairs in both native Perso‑Arabic script and Romanized Urdu. It evaluates multiple tasks—translation, paraphrasing, idiom span detection, and back‑translation—using various prompting strategies, and finds that state‑of‑the‑art large language models outperform traditional neural machine translation systems, especially in preserving figurative meaning. The study also highlights challenges posed by the lack of standardized orthography in Romanized Urdu, which affects consistency and idiom span detection.
By Muhammad Farmal Khan, Mousumi Akter
arXiv:2606. 07167v1 Announce Type: cross Abstract: Meaningful multilingual evaluation must test models in the target language and educational context.
By Ahmer Tabassum, Sarfraz Ahmad, Hasan Iqbal, Owais Aijaz, Momina Ahsan, Preslav Nakov
Multilingual Language Models like mBERT are widely used for low-resource NLP, yet their adaptation to morphologically inconsistent languages such as Roman Urdu remains underexplored. Roman Urdu spelling variation causes severe sub-word fragmentation, averaging 1.
The paper introduces UrduFactBench and UrduFactQA, two hand‑annotated benchmarks for claim verification and factual consistency evaluation in Urdu, created through a multi‑stage annotation process with native speakers. It also presents UrduFactCheck, a modular fact‑checking framework that uses both monolingual and translation‑based evidence retrieval to address the scarcity of high‑quality Urdu evidence. Experiments on twelve LLMs show that translation‑augmented pipelines outperform monolingual ones, highlighting ongoing challenges for open‑source models in Urdu.
By Sarfraz Ahmad, Hasan Iqbal, Momina Ahsan, Numaan Naeem, Muhammad Ahsan Riaz Khan, Arham Riaz, Muhammad Arslan Manzoor, Yuxia Wang, Preslav Nakov
The paper investigates how multilingual large language models perform when generating stories in Urdu, a low‑resource language. The authors created a corpus of 93 Urdu stories produced by GPT‑5.1, Qwen‑3‑Max, and DeepSeek‑3.1, and manually annotated errors across a nine‑label taxonomy covering linguistic, semantic, and cultural aspects. Findings reveal frequent grammatical and semantic mistakes, lack of coherence, unnatural repetition, and pervasive cultural shallowness, with few‑shot prompting failing to resolve many of these issues.
By Farah Adeeba, Abdul Rafae Khan, Rajesh Bhatt, Hassan Sajjad
The paper presents the Multilingual FrameNet Corpus (mFNC), a resource that expands the English Berkeley FrameNet by integrating and harmonizing language‑specific corpora in nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian, and Swedish. Experiments with various model architectures on mFNC consistently surpass existing state‑of‑the‑art Frame Semantic Parsers in both multilingual and cross‑lingual scenarios, highlighting the value of multilingual training data. The mFNC and the trained Frame Semantic Parser models are publicly released on GitHub.
By Beatrice Fiuman\`o, Nicolas Lazzari, Simone Paolo Ponzetto, Valentina Presutti
arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv:2607. 08063v1 Announce Type: cross Abstract: Unsupervised constituency parsing aims to accurately induce latent tree structures from raw text alone.
By Ryosuke Yamaki, Daichi Mochihashi, Nobutaka Shimada, Tadahiro Taniguchi
arXiv:2609.13721v1 Announce Type: new
Abstract: Quantum Natural Language Processing (QNLP) uses pregroup grammars to translate grammatical structure into diagrammatic representations and quantum circ...
By Gautami Sanjay Naik, Krishna Bhatia, Mithun Paul Saint-Germain, H Aswath Babu
arXiv:2011.03783v3 Announce Type: replace-cross
Abstract: In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with an...
By Lifeng Han, Najet Hadj Mohamed, Malak Rassem, Gareth Jones, Alan Smeaton, Goran Nenadic
ThaiTrees is a 342‑million‑token corpus of Thai text spanning news, Wikipedia, spoken transcripts, and social media, automatically parsed under the Universal Dependencies framework. The authors provide a reproducible pipeline for cleaning, processing, and parsing the data, and release the resulting CoNLL‑U files and a frequency lexicon in machine‑readable formats. This resource enables researchers to search grammatical relations and study syntactic distributions at scale.
By Attapol T. Rutherford, Papatchol Thientong