Llama 3.1 - 405B, 70B & 8B with multilinguality and long context
Related stories
SinLlama -- A Large Language Model for Sinhala
The paper introduces SinLlama, the first decoder‑based open‑source large language model with explicit support for Sinhala. By extending Llama‑3‑8B, adding Sinhala‑specific tokenizer vocabulary, and performing continual pre‑training on a cleaned 10‑million‑token Sinhala corpus, the authors created a model that surpasses both the base and instruction‑fine‑tuned variants of Llama‑3‑8B on three text classification tasks. This work addresses the underrepresentation of low‑resource languages in open‑source LLMs.
Tracing Stereotypes from Representation to Output in Multilingual LLMs
The paper investigates how multilingual large language models (LLMs) encode and express stereotypes across different languages. By applying linear probing, attribution patching, sparse autoencoders (SAEs), and feature ablation to Llama‑3.1‑8B, Qwen3‑8B, and Gemma‑2‑9B, the authors find that probe performance peaks much earlier than attribution, indicating a separation of 36‑53% of model depth. They observe that only a small fraction (6‑18%) of residual‑stream features exhibit language‑agnostic effects, and none are category‑agnostic, highlighting the need to measure decodability, output influence, and cross‑lingual ablation effects separately.
StackLLaMA: A hands-on guide to train LLaMA with RLHF
Tracing Stereotypes from Representation to Output in Multilingual LLMs
arXiv:2609.08322v1 Announce Type: cross Abstract: Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is...
5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs
5-Dialects-BN is a new Bangla dialect benchmark that aligns Romanized transliteration with dialectal text, Standard Bangla, English, and subjectivity labels across five regional varieties. The dataset contains 6,000 manually annotated entries from Chittagong, Barisal, Noakhali, Sylhet, and Rangpur, each enriched with five aligned annotations produced and cross‑validated by native speakers and linguistics students. It supports tasks such as dialect identification, normalization, translation, subjectivity classification, and efficient fine‑tuning of multilingual LLMs.
Script Fragmentation and Format: What Drives the English-Bengali Performance Gap in Open LLMs?
arXiv:2507.23248v2 Announce Type: replace-cross Abstract: Bengali is spoken by more than 230 million people, yet no standardized instrument evaluates large language models (LLMs) on Bengali across th...
Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM
Code Llama: Llama 2 learns to code
Introducing The World's Largest Open Multilingual Language Model: BLOOM
Multilinguality of Large Language Models From a Structural Perspective
arXiv:2606. 01800v1 Announce Type: cross Abstract: Large language models (LLMs) have excelled in processing multiple languages through pre- and post-training on multilingual data, even though English dominates the training data.