arXiv AI

Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study

arXiv Computation and Language
2d ago

COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages

COILD is an Indic‑centric parallel corpus that contains over 1.16 million human‑translated and verified sentence pairs across 20 Indian language pairs from four language families. The corpus is sourced from original Indian language materials in eight domains, and a 2,000‑sentence domain‑centric benchmark is provided for consistent multilingual evaluation. Experiments with IndicTrans2‑Distilled and NLLB‑200 show consistent improvements in automatic metrics and human judgments, underscoring the value of high‑quality Indic‑centric data.

By Kshetrimayum Boynao Singh, Nitin Kumar Mishra, Palash Pratim Dutta, Atai Waris Khan, Aparna Kaushik, Avinash Kumar, Deeksha, Deepak Kumar, Saroj Kumar Jha, Saloka Sengupta, Anansa Roy, Umalatha Kannoth, Saifulla Samar, Meena Sharma, Manpreet Kaur, Jyoti Sharma, Ashwini Vaidya, Muralikrishna SN, Md Shad Akhtar, Poonam Bansal, Amita Dev, Sanasam Ranbir Singh, Samit Bhattacharya, Tanmoy Chakraborty, Asif Ekbal
arXiv AI
Sep 3

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.

By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
arXiv Machine Learning
Jul 28

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.

By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv Computation and Language
Sep 18

Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali

Viveka-Insight is a bilingual resource and open‑source pipeline for Swami Vivekananda’s complete works, providing a structure‑preserving parse of 32,694 paragraphs and 168,842 sentences, a cross‑lingual concept graph with 8,362 language‑agnostic concepts, a bilingual alias inventory of 60,850 surface forms, and a human‑annotated set of 200 paragraph‑concept edges. The resource enables citation‑grounded retrieval across the English and Bengali corpora, achieving Recall@10 of 0.86 for known‑item cross‑lingual queries and demonstrating concept‑extraction precision of 0.60 (up to 0.71 with confidence filtering). The design is intended to be transferable to other multilingual classical corpora.

By Tamal Maharaj
arXiv Machine Learning
Jul 24

Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models

arXiv:2607. 20442v1 Announce Type: cross Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in July 2022 across two categories (Economy and IT/Science; 77/23 split), with train/validation/test partitions of 22,194 / 2,466 / 2,740 and a mean per-record document-to-summary character-compression ratio of 6.

By Daekeun Kim
arXiv Computation and Language
Aug 27

Padamitra: Grounded Glossary Generation for Classical Sanskrit

The paper introduces grounded glossary generation, a structured NLP task that asks models to recover semantically meaningful Sanskrit phrases and provide translation‑grounded meanings from a sloka‑translation pair, mirroring the traditional patha commentary practice. A benchmark of 31,316 sloka‑translation‑glossary triples from the Valmiki Ramayana and Srimad Bhagavatam is built, evaluated with Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Experiments with Gemma‑3n‑E4B, Gemma‑3‑12B, Phi‑4, and Qwen3.5‑9B show that instruction fine‑tuning outperforms prompting, and explicit segmentation further improves results, though over‑segmentation of sandhi and samasa compounds remains the main error source, highlighting morphological modeling as a key bottleneck.

By Manoj Balaji Jagadeeshan, Sai Pragnaan Marala, Pawan Goyal
arXiv AI
Sep 3

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

VakyArth is the first pragmatic benchmark for Indic languages, covering Hindi, Punjabi, Tamil, and Malayalam. It tests models on five pragmatic phenomena—deixis, speech acts, implicature, social pragmatics, and coherence—using multiple-choice questions, natural language inference, and translation tasks authored by native speakers. Evaluation of multilingual LLMs shows consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions, with systematic differences across languages and tasks.

By Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda, Anand Madhusoodanan, Geyang Guo, Wei Xu, Junyi Jessy L