Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages.
The paper introduces MMSAFE, a multi-layer framework designed to identify safety-degrading data in multilingual large language models. It shows that safety signals are distributed across multiple layers and only partially shared across languages, unlike the single-layer assumption used in monolingual settings. Experiments demonstrate that MMSAFE reduces harmful-response rates by 60% compared to random filtering and outperforms the best single-layer baseline across various models, languages, and safety benchmarks.
By Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan
arXiv:2606. 08451v1 Announce Type: cross Abstract: Safety-aligned large language models often exhibit sycophancy, which is the tendency to affirm users' opinions regardless of factual accuracy.
By Arya Shah, Himanshu Beniwal, Mayank Singh, Chaklam Silpasuwanchai
arXiv:2606. 28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task.
By Will Hawkins, Kaivalya Rawal, Jonathan Rystr{\o}m, Stratis Tsirtsis, Zihao Fu, Greta Warren, Ryan Brown, Eoin Delaney, Sandra Wachter, Brent Mittelstadt, Chris Russell
arXiv:2608.22490v1 Announce Type: new
Abstract: Safety alignment helps models adhere to human values, but it often reduces response utility. We ask a critical but understudied question: Does safety a...
By Chanwoong Yoon, Jungsoo Park, Alan Ritter
arXiv:2607. 10112v1 Announce Type: cross Abstract: Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings.
By Chigozirim Ifebi, Brent Kong, Ayushi Mehrotra