arXiv AI By Ahmer Tabassum, Sarfraz Ahmad, Hasan Iqbal, Owais Aijaz, Momina Ahsan, Preslav Nakov

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

Read the original on arXiv AI →

arXiv:2606. 07167v1 Announce Type: cross Abstract: Meaningful multilingual evaluation must test models in the target language and educational context.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 14

UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking

The paper introduces UrduFactBench and UrduFactQA, two hand‑annotated benchmarks for claim verification and factual consistency evaluation in Urdu, created through a multi‑stage annotation process with native speakers. It also presents UrduFactCheck, a modular fact‑checking framework that uses both monolingual and translation‑based evidence retrieval to address the scarcity of high‑quality Urdu evidence. Experiments on twelve LLMs show that translation‑augmented pipelines outperform monolingual ones, highlighting ongoing challenges for open‑source models in Urdu.

By Sarfraz Ahmad, Hasan Iqbal, Momina Ahsan, Numaan Naeem, Muhammad Ahsan Riaz Khan, Arham Riaz, Muhammad Arslan Manzoor, Yuxia Wang, Preslav Nakov
arXiv Computation and Language
Sep 4

Evaluating Large Language Models on Urdu Idioms

The paper introduces a new benchmark for Urdu‑to‑English idiomatic translation, featuring 4,000 manually verified sentence pairs in both native Perso‑Arabic script and Romanized Urdu. It evaluates multiple tasks—translation, paraphrasing, idiom span detection, and back‑translation—using various prompting strategies, and finds that state‑of‑the‑art large language models outperform traditional neural machine translation systems, especially in preserving figurative meaning. The study also highlights challenges posed by the lack of standardized orthography in Romanized Urdu, which affects consistency and idiom span detection.

By Muhammad Farmal Khan, Mousumi Akter
arXiv Machine Learning
Sep 11

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

The paper investigates how multilingual large language models perform when generating stories in Urdu, a low‑resource language. The authors created a corpus of 93 Urdu stories produced by GPT‑5.1, Qwen‑3‑Max, and DeepSeek‑3.1, and manually annotated errors across a nine‑label taxonomy covering linguistic, semantic, and cultural aspects. Findings reveal frequent grammatical and semantic mistakes, lack of coherence, unnatural repetition, and pervasive cultural shallowness, with few‑shot prompting failing to resolve many of these issues.

By Farah Adeeba, Abdul Rafae Khan, Rajesh Bhatt, Hassan Sajjad
arXiv AI
Aug 26

'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection

The paper "Ghaib in Translation" investigates how large language models (LLMs) handle Urdu, a widely spoken language that is largely absent from safety evaluations. Five prominent LLMs—GPT‑4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen‑2.5, and Llama‑3.1—were tested on six datasets covering Nastaliq Urdu, Roman Urdu, English, and code‑switched Urdu‑English. The study found significant label instability between original‑script and English‑translation classifications, with missed‑in‑Urdu rates ranging from 2.4% to 9.9% (median 4.3%). A review of 205 papers across nine ALW/WOAH editions revealed no dedicated Urdu research, underscoring the language’s neglect in current safety research.

By Fawzia Zehra (Fuzzy), Kara-Isitt, Sonal Khosla, Stephen Swift