arXiv AI

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

arXiv:2608. 14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages.

arXiv Machine Learning
Sep 22

Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs

The paper introduces MMSAFE, a multi-layer framework designed to identify safety-degrading data in multilingual large language models. It shows that safety signals are distributed across multiple layers and only partially shared across languages, unlike the single-layer assumption used in monolingual settings. Experiments demonstrate that MMSAFE reduces harmful-response rates by 60% compared to random filtering and outperforms the best single-layer baseline across various models, languages, and safety benchmarks.

By Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan
arXiv AI
Aug 20

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

The paper titled "Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs" highlights that current safety alignment training for large language models is predominantly English-centric, leading to failures in non‑English languages. It introduces INCLUDE, a multilingual benchmark with 2,604 prompts in six languages (English, Hindi, Bengali, Marathi, Tamil, and Hinglish) to measure Indian‑centric socio‑cultural biases. Evaluation of ten open‑ and closed‑source LLMs shows that Bengali models exhibit the highest bias scores among open‑source models, while English shows the lowest bias in open‑source but the highest in closed‑source models.

By Namya Bhatnagar
arXiv AI
Jun 16

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.

By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang