arXiv Machine Learning By Erik Arakelyan, Khatun Avetisyan, Meri Davtyan, Heghine Grigoryan, Nane Khachatryan, Hayk Shahsuvaryan, Henrik Sergoyan, Vahan Martirosyan

From Zero to Hero: An Open LLM Ecosystem for Armenian

Read the original on arXiv Machine Learning →

The paper introduces the first open Armenian large language model, arm‑gemma‑e4b, trained on two newly released datasets: ArmWeb, a 4.37 million‑document news corpus, and ArmSTEM, a 373 k English‑Armenian math and science problem set with verified step‑by‑step solutions. Continued pretraining of Gemma‑4‑E4B on these datasets outperforms all existing open Armenian models and demonstrates that adding a small portion of verified translated STEM data can restore knowledge lost during news‑only pretraining. The authors also reveal significant overlap between major public Armenian corpora and web‑derived evaluation panels, and they provide all data, models, and code openly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 25

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.

By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
arXiv Computation and Language
Sep 10

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

arXiv:2607.00890v2 Announce Type: replace Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synth...

By Maximilian Idahl, J\"org Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, Andr\'e F. T. Martins, Jenna Kanerva, Filip Ginter, Matthias Lindemann, Tim Isbister, Birger Moell, Jonas Lindh, Jan Haji\v{c}, Jenia Jitsev, Andrey Kutuzov, Stephan Oepen, Gema Ram\'irez-S\'anchez
arXiv Computation and Language
2d ago

ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch

The paper introduces ufakzeka-1, a 151‑million‑parameter Turkish language model trained from scratch on 13.5 B tokens. It details the tokenizer, a three‑stage pretraining schedule, post‑training data augmentation, and a comprehensive evaluation suite that includes release gates, a rule‑checked conversation sweep, and hand tests—all with prompts excluded from training data. The authors report three key findings about safety gate performance, training‑seed variance, and data‑round effects, and they release the model weights, data recipe, evaluation code, and spend ledger under Apache‑2.0.

By Sait Furkan Teke (ufak AI)
arXiv AI
Jun 12

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.

By Marek \v{S}uppa, Andrej Ridzik, Daniel Hl\'adek, Nat\'alia K\v{n}a\v{z}ekov\'a, Vikt\'oria Ondrejov\'a