The paper introduces SinLlama, the first decoder‑based open‑source large language model with explicit support for Sinhala. By extending Llama‑3‑8B, adding Sinhala‑specific tokenizer vocabulary, and performing continual pre‑training on a cleaned 10‑million‑token Sinhala corpus, the authors created a model that surpasses both the base and instruction‑fine‑tuned variants of Llama‑3‑8B on three text classification tasks. This work addresses the underrepresentation of low‑resource languages in open‑source LLMs.
By H. W. K. Aravinda, Rashad Sirajudeen, Samith Karunathilake, Nisansa de Silva, Surangika Ranathunga, Rishemjit Kaur
arXiv:2012. 02110v2 Announce Type: replace-cross Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimized version, RoBERTa.
By Raphael Scheible, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer, Martin Boeker
NE‑BERT is a multilingual encoder trained on about 8.3 million sentences from nine Northeast Indian languages plus Hindi and English. Using weighted sampling and a custom SentencePiece tokenizer, it achieves significantly lower perplexity than IndicBERT‑V2, MuRIL, and mBERT, and improves tokenization fertility. The model also addresses vocabulary fragmentation in extremely low‑resource languages through aggressive upsampling, and its effectiveness is validated on part‑of‑speech tagging for three of the languages.
By Badal Nyalang
arXiv:2606. 20089v1 Announce Type: cross Abstract: Persian pretrained language models (PLMs) are still limited by the scarcity of large-scale, high-quality pretraining corpora and by insufficient evaluation beyond standard classification and NER tasks.
By Arash Ghafouri, Mahdi Firouzmandi, Hossein Saberi, Mohammad Reza Hasani Ahangar
arXiv:2608. 14649v1 Announce Type: new Abstract: We present dLLM-SetScore, a training-free method that uses discrete masked-diffusion language models for multi-label text classification.
By Pawan Kumar
arXiv:2506. 15138v2 Announce Type: replace-cross Abstract: Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length and generation cost.
By Gyeongje Cho, Yeonkyoung So, Sangmin Lee, Jaejin Lee
arXiv:2606. 24825v1 Announce Type: cross Abstract: Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing.
By Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Raviraj Joshi
Multilingual Language Models like mBERT are widely used for low-resource NLP, yet their adaptation to morphologically inconsistent languages such as Roman Urdu remains underexplored. Roman Urdu spelling variation causes severe sub-word fragmentation, averaging 1.
BanglaMamba explores Mamba-based State Space Models (SSMs) as a computationally efficient alternative for Bangla fake news detection. Compared to BanglaBERT and a custom BERT trained from scratch, BanglaMamba achieves a Macro‑F1 score of 0.9029, close to the 0.9057 of the custom BERT, while delivering 2.2× higher inference throughput and 49% lower peak GPU memory usage. Cross‑dataset evaluation shows BanglaBERT generalizes better, underscoring the value of large‑scale pretraining.
By M. K. Khalidi Siam
IndicDetect is a benchmark for evaluating AI‑generated text detection in Hindi, Telugu, and Tamil. It pairs curated human‑written texts with LLM‑generated counterparts across multiple domains and generators, testing detectors under domain shift, generator shift, and adversarial perturbation. The study shows that supervised neural detectors fail to generalize to unseen generators and attacks, with Hindi experiencing the greatest degradation, indicating that robustness—not peak accuracy—is the main weakness in Indic language detectors.
By Bhaskar Ganesh Devalla, Junchao Wu, Nilesh Dokuparthi, Greeshma Yaluru, Tatiana Muniz Rodriguez, Lidia S. Chao, Derek F. Wong
arXiv:2606. 03334v1 Announce Type: cross Abstract: Our submission presented in this paper is for SemEval-2026 Task 9: Multilingual Text Classification Challenge - Polarization Detection and it covers all three subtasks: (1) binary polarization detection, (2) polarization type classification and (3) polarization manifestation identification.
By Pritam Kadasi, Anuj Tiwari, Mayank Singh
arXiv:2606. 29614v1 Announce Type: cross Abstract: This study examines whether supervised fine-tuning remains necessary for Turkish sentiment analysis in the era of large language models.
By Sercan Karaka\c{s}, Yusuf \c{S}im\c{s}ek