arXiv Computation and Language
4d ago

SinLlama -- A Large Language Model for Sinhala

The paper introduces SinLlama, the first decoder‑based open‑source large language model with explicit support for Sinhala. By extending Llama‑3‑8B, adding Sinhala‑specific tokenizer vocabulary, and performing continual pre‑training on a cleaned 10‑million‑token Sinhala corpus, the authors created a model that surpasses both the base and instruction‑fine‑tuned variants of Llama‑3‑8B on three text classification tasks. This work addresses the underrepresentation of low‑resource languages in open‑source LLMs.

By H. W. K. Aravinda, Rashad Sirajudeen, Samith Karunathilake, Nisansa de Silva, Surangika Ranathunga, Rishemjit Kaur
arXiv Machine Learning
Jun 2

GottBERT: a pure German Language Model

arXiv:2012. 02110v2 Announce Type: replace-cross Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimized version, RoBERTa.

By Raphael Scheible, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer, Martin Boeker
arXiv AI
Aug 20

NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages

NE‑BERT is a multilingual encoder trained on about 8.3 million sentences from nine Northeast Indian languages plus Hindi and English. Using weighted sampling and a custom SentencePiece tokenizer, it achieves significantly lower perplexity than IndicBERT‑V2, MuRIL, and mBERT, and improves tokenization fertility. The model also addresses vocabulary fragmentation in extremely low‑resource languages through aggressive upsampling, and its effectiveness is validated on part‑of‑speech tagging for three of the languages.

By Badal Nyalang