arXiv AI

Rethinking Indic AI from a Lens of Cultural Heritage Preservation

arXiv:2607. 06544v1 Announce Type: new Abstract: As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant interest in studying how AI impacts the linguistic and cultural foundations of this civilization.

Hugging Face Trending Papers
Jun 20

Plurification in/of language technology -- The integration of culture in next-generation AI

The paper explores how "culture" can be operationalised in Natural Language Processing (NLP) and what this reveals about the possibilities and limits of considering a plurality of cultural backgrounds in technological design. It proposes that cultural alignment cannot be achieved only by adding more examples of "other cultures", rather it requires plural epistemologies: allowing multiple, locally grounded ways of knowing.

arXiv Computation and Language
Aug 28

Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

The study investigates how large language models (LLMs) handle diverse Indian oral traditions, using the Rajasthani Pabuji epic, Tamil Sangam poetry, and Bengali folk tales as case studies. By prompting Claude Sonnet and Gemini with 54 generation requests across generic, culturally specific, and regional-language prompts, the authors measured reference drift and cross-tradition convergence using Sentence‑BERT embeddings. Results show that while outputs stay closer to their own tradition than to others, there is significant cross‑tradition similarity (0.52–0.66), indicating partial homogenisation; moreover, regional‑language prompting consistently reduced fidelity to authentic traditions.

By Paarth Singh Rathore
arXiv AI
Sep 3

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

VakyArth is the first pragmatic benchmark for Indic languages, covering Hindi, Punjabi, Tamil, and Malayalam. It tests models on five pragmatic phenomena—deixis, speech acts, implicature, social pragmatics, and coherence—using multiple-choice questions, natural language inference, and translation tasks authored by native speakers. Evaluation of multilingual LLMs shows consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions, with systematic differences across languages and tasks.

By Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda, Anand Madhusoodanan, Geyang Guo, Wei Xu, Junyi Jessy L
arXiv Machine Learning
Jul 28

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.

By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv AI
6d ago

NaijaNLP: A Survey of Nigerian Low-Resource Languages

The paper surveys NLP research on Nigeria’s three major low‑resource languages—Hausa, Yoruba, and Igbo—covering over 500 languages spoken by 175 million people. It reviews 293 studies, finding that only 27.6% produced new linguistic resources, indicating a heavy reliance on repurposing existing data. The authors highlight under‑explored challenges such as morphological analysis and diacritic representation, and call for collaborative resource enrichment and community support to advance NaijaNLP and low‑resource NLP more broadly.

By Isa Inuwa-Dutse
Hugging Face Trending Papers
Aug 10

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked.

arXiv Computation and Language
Sep 3

MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

MemeCULT-1K is a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, each paired with a cultural context note and three human-written explanations, plus an additional set of 54 Bengali regional dialect memes. The study evaluates thirteen vision‑language models under meme‑only and context‑aware settings, showing that providing minimal cultural context consistently improves performance across all models and languages. Error analysis indicates closed‑source models struggle with entity and reference misidentification, while open‑source models are limited by broader cultural knowledge gaps, especially in linguistic and phonological aspects.

By Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir, Mueeze Al Mushabbir, Mohammed Saidul Islam, Mir Rayat Imtiaz Hossain, Md Tahmid Rahman Laskar, Sabbir Ahmed