arXiv:2606. 09178v1 Announce Type: cross Abstract: Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target languages - an approach that converts surface-level linguistic form while failing to reflect the cultural context embedded in threat scenarios, social norms, and legal frameworks.
By Hyeji Choi, Yongtaek Lim, Minwoo Kim
Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target languages - an approach that converts surface-level linguistic form while failing to reflect the cultural context embedded in threat scenarios, social norms, and legal frameworks. We construct paired DT and culturally-adapted (CA) datasets via 1:1 seed matching for four languages - Korean (KO), Japanese (JA), Thai (TH), and Khmer (KM) - and compare Attack Success Rate (ASR) and Cultural Realism scores across four open-source LLM.
arXiv:2605.28013v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vi...
By Yongwoo Kim, Sojung An, Yunjin Park, Jungwon Yoon, Dujin Lee, HyunBeom Cho, Jaewon Lee, Wonhyuk Lee, Youngchol Kim, JeongYeop Kim, Donghyun Kim
arXiv:2605. 17173v2 Announce Type: replace-cross Abstract: Large language models exhibit safety degradation in non-English languages.
By Max Zhang, Ameen Patel, Sang T. Truong, Sanmi Koyejo
arXiv:2606. 28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task.
By Will Hawkins, Kaivalya Rawal, Jonathan Rystr{\o}m, Stratis Tsirtsis, Zihao Fu, Greta Warren, Ryan Brown, Eoin Delaney, Sandra Wachter, Brent Mittelstadt, Chris Russell
Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages.
arXiv:2606. 01322v1 Announce Type: cross Abstract: Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, critically underexplored.
By Victor Akinode, Senyu Li, Wassim Hamidouche, Waqas Zamir, Inbal Becker-Reshef, David Ifeoluwa Adelani
arXiv:2607. 02079v1 Announce Type: cross Abstract: We present HaloGuard 1.
By Navaneeth Sangameswaran, Preetham S, Ashmiya Lenin
arXiv:2607. 01859v1 Announce Type: new Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching.
By Joshua Adrian Cahyono
Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.
IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark includes 7,200 adversarial prompts covering ten safety-critical content categories, six persuasive strategies, and four languages (Hindi, Bengali, Marathi, Punjabi). Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, revealing that current English-centric safety evaluations miss important multilingual vulnerabilities.
By Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal
IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark covers ten safety-critical content categories, six persuasive strategies, and four languages—Hindi, Bengali, Marathi, and Punjabi—producing 7,200 adversarial prompts. Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, highlighting gaps in current English-centric safety assessments.