arXiv AI

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

arXiv:2606. 28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task.

arXiv AI
Aug 18

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

arXiv:2608. 14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages.

By Valdini Douglace Lemofouet, Blessing Ngozi Uzor, Paula Chikaodinaka Anyanwu, Danielle Blanche Kapsa, Sukairaj Hafiz Imam, P Sam Sahil, Abigail Oppong, Tassallah Abdullahi, Clemencia Siro, Idris Abdulmumin, Seid Muhie Yimam, Shamsuddeen Hassan Muhammad
arXiv Machine Learning
Sep 22

Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs

The paper introduces MMSAFE, a multi-layer framework designed to identify safety-degrading data in multilingual large language models. It shows that safety signals are distributed across multiple layers and only partially shared across languages, unlike the single-layer assumption used in monolingual settings. Experiments demonstrate that MMSAFE reduces harmful-response rates by 60% compared to random filtering and outperforms the best single-layer baseline across various models, languages, and safety benchmarks.

By Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan
Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.

arXiv AI
Aug 20

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

The paper titled "Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs" highlights that current safety alignment training for large language models is predominantly English-centric, leading to failures in non‑English languages. It introduces INCLUDE, a multilingual benchmark with 2,604 prompts in six languages (English, Hindi, Bengali, Marathi, Tamil, and Hinglish) to measure Indian‑centric socio‑cultural biases. Evaluation of ten open‑ and closed‑source LLMs shows that Bengali models exhibit the highest bias scores among open‑source models, while English shows the lowest bias in open‑source but the highest in closed‑source models.

By Namya Bhatnagar
arXiv Machine Learning
Aug 5

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.

By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y