Hugging Face Trending Papers

Three Buddhist Vocabularies: Computational Stylometry of the English Pali Canon across Sutta, Vinaya, and Abhidhamma

Read the original on Hugging Face Trending Papers →

We present a computational stylometric analysis of the Tipitaka across all three Pitakas in English translation, extending earlier work on the Sutta Pitaka alone. The corpus spans 134,831 segments from Bhikkhu Sujato's Sutta Pitaka (114,591 segments, CC0), Bhikkhu Brahmali's Vinaya Pitaka (7,923 segments, CC0 2026), I.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 21

TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar

TatBLiMP is the first benchmark of linguistic minimal pairs for the Tatar language, covering 16 morphosyntactic phenomena across 1,248 sentence pairs that differ by a single morpheme. Each pair contains one grammatical and one ungrammatical sentence, with the ungrammatical version generated by a deterministic perturbation and ratified by a native speaker. The benchmark evaluates models by comparing their assigned probabilities, allowing assessment without text generation or parsing, and tracks performance across from-scratch, cross‑lingual, and multilingual large language models.

By Ilshat Saetov, Dmitry Gaynullin
arXiv Computation and Language
5d ago

SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0

arXiv:2603.10861v2 Announce Type: replace Abstract: SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication date...

By Nevidu Jayatilleke, Nisansa de Silva, Uthpala Nimanthi, Gagani Kulathilaka, Azra Safrullah, Johan Sofalas
arXiv Computation and Language
Sep 18

Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali

Viveka-Insight is a bilingual resource and open‑source pipeline for Swami Vivekananda’s complete works, providing a structure‑preserving parse of 32,694 paragraphs and 168,842 sentences, a cross‑lingual concept graph with 8,362 language‑agnostic concepts, a bilingual alias inventory of 60,850 surface forms, and a human‑annotated set of 200 paragraph‑concept edges. The resource enables citation‑grounded retrieval across the English and Bengali corpora, achieving Recall@10 of 0.86 for known‑item cross‑lingual queries and demonstrating concept‑extraction precision of 0.60 (up to 0.71 with confidence filtering). The design is intended to be transferable to other multilingual classical corpora.

By Tamal Maharaj
arXiv Machine Learning
Jul 28

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

arXiv:2607. 23322v1 Announce Type: cross Abstract: Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains.

By Shwetha Singaravelu, Gayathri Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani