arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).
By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
arXiv:2608.30092v1 Announce Type: cross
Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single...
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
arXiv:2608.28611v1 Announce Type: cross
Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trai...
By Isha Narang, Sneh Gosai, Mayank Singh
arXiv:2609.14829v1 Announce Type: cross
Abstract: We introduce Enemray, a Hassaniya-centric language model that enables general-purpose interaction in Hassaniya. Enemray is trained around a stability...
By Cheikh Ahmed
The paper investigates how to choose language models (teachers) for generating multilingual synthetic data used to fine‑tune smaller student models. By evaluating 10 teacher models across six diverse languages and training 240 students, the authors find that teacher effectiveness is not driven by model size but by data qualities such as prompt diversity, length, and fluency, which explain most of the variance in student performance. Practical guidelines are offered, including matching teacher and student families and using translated prompts to improve outcomes for low‑resource languages.
By Lester James V. Miranda, Ivan Vuli\'c, Anna Korhonen
arXiv:2605.29637v2 Announce Type: replace
Abstract: Large language models often exhibit a substantial gap between their performance in English and in lower-resourced languages on equivalent knowledge...
By Debajyoti Mazumder, Divyansh Pathak, Prashant Kodali, Aditya Joshi, Akshay Agarwal, Jasabanta Patro