arXiv:2608.30609v1 Announce Type: cross
Abstract: Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to addres...
By Lukas Borggren, Jenny Kunz, Marco Kuhlmann
The paper introduces the problem of cross‑lingual loopholes in large language model (LLM) unlearning, where forgetting a fact in one language can leave it accessible in others. It presents a new 174‑language benchmark, the Cross‑Lingual Unlearning Tensor, and proposes COVER, a method that selects a subset of source languages to maximize unlearning coverage under a language budget. Experiments show COVER reduces residual knowledge by 7.8–27.3% compared to uniform selection and works on both synthetic and real low‑resource news data.
By Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav
arXiv:2609.37076v1 Announce Type: new
Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this is...
By Puning Yang, Qizhou Wang, Junchi Yu, Bo Han, Xiuying Chen
We introduce Dango, a 1. 8B-parameter large language model designed for controlled studies of L1-to-L2 (Japanese-to-English) transfer in second language acquisition (SLA).
arXiv:2510. 15551v2 Announce Type: replace-cross Abstract: Any piece of knowledge is usually expressed in one or a handful of natural languages on the web or in any large corpus.
By Vihari Piratla, Purvam Jain, Darshan Singh, Trevor Cohn, Preethi Jyothi, Partha Talukdar
Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results....
arXiv:2606.12234v2 Announce Type: replace
Abstract: Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the invol...
By Iuri Macocco, Pau Rodr\'iguez, Arno Blaas, Luca Zappella, Marco Baroni, Xavier Suau
arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.
By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin
LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.
By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
The paper compares Knowledge-Graph Based Augmentation (Graph-RAG) with Retrieval-Augmented Generation (RAG) for answering culturally specific questions. Using the LatamQA dataset, Graph-RAG, built automatically from Wikipedia via KGGen, matches RAG performance and reduces the base LLM’s error by 72% with a standard KG and 78% with a benchmark-aware variant. The approach also transfers zero‑shot to Portuguese, showing multilingual applicability.
By Pablo Poulenard, Yannis Karmim, Valentin Barri\`ere
arXiv:2607. 19243v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages.
By Alexander Manev
arXiv:2410. 07809v2 Announce Type: replace-cross Abstract: Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost.
By G\"urkan Soykan, G\"ozde G\"ul \c{S}ahin