Rethinking Cross-lingual Gaps from a Statistical Viewpoint
arXiv:2510. 15551v2 Announce Type: replace-cross Abstract: Any piece of knowledge is usually expressed in one or a handful of natural languages on the web or in any large corpus.
arXiv:2510. 15551v2 Announce Type: replace-cross Abstract: Any piece of knowledge is usually expressed in one or a handful of natural languages on the web or in any large corpus.
arXiv:2510.27183v3 Announce Type: replace Abstract: The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. How...
arXiv:2609.37882v1 Announce Type: cross Abstract: Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other Afri...
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions.
The paper introduces the problem of cross‑lingual loopholes in large language model (LLM) unlearning, where forgetting a fact in one language can leave it accessible in others. It presents a new 174‑language benchmark, the Cross‑Lingual Unlearning Tensor, and proposes COVER, a method that selects a subset of source languages to maximize unlearning coverage under a language budget. Experiments show COVER reduces residual knowledge by 7.8–27.3% compared to uniform selection and works on both synthetic and real low‑resource news data.
arXiv:2608.20362v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an ans...
arXiv:2608.30462v1 Announce Type: cross Abstract: Large language models exhibit substantial performance variation across languages, even when solving semantically equivalent tasks. Existing analyses...
The paper introduces CLLPU, a multilingual benchmark for evaluating how well large language models can unlearn specific knowledge while controlling its propagation across languages. CLLPU defines two forgetting scenarios—common-goal forgetting, which requires suppression across all languages, and language-conditioned forgetting, which limits suppression to a single language. Using 800 knowledge-unit pairs and 72,000 QA instances in ten languages, the authors test six methods on Llama‑3.1‑8B‑Instruct and find that universal suppression often fails, while language‑specific suppression can unintentionally spread to other languages, highlighting the difficulty of propagation control in multilingual unlearning.
arXiv:2607. 04814v1 Announce Type: cross Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale.
The paper investigates whether multilingual language models transfer factual knowledge from one language to another during continued pretraining. Using an English-pretrained model continued on Persian data with systematically removed facts, the authors create SIFT, a dataset of 500 triples across 20 relations, split by cultural origin. Their findings indicate that factual transfer is minimal, especially for Persian-related facts, and that simple removal strategies or easy negative candidate sets can overestimate transfer.
arXiv:2609.37543v1 Announce Type: new Abstract: Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-reso...
arXiv:2608. 13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users.