arXiv Machine Learning By Shwetha Singaravelu, Gayathri Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

Read the original on arXiv Machine Learning →

arXiv:2607. 23322v1 Announce Type: cross Abstract: Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 18

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).

By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
arXiv Computation and Language
Sep 1

Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation

The paper investigates how to choose language models (teachers) for generating multilingual synthetic data used to fine‑tune smaller student models. By evaluating 10 teacher models across six diverse languages and training 240 students, the authors find that teacher effectiveness is not driven by model size but by data qualities such as prompt diversity, length, and fluency, which explain most of the variance in student performance. Practical guidelines are offered, including matching teacher and student families and using translated prompts to improve outcomes for low‑resource languages.

By Lester James V. Miranda, Ivan Vuli\'c, Anna Korhonen