arXiv AI

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection

arXiv:2606. 19168v1 Announce Type: new Abstract: To achieve deeper safety alignment for large language models (LLMs), recent efforts have studied how to push safety interventions earlier into the pretraining stage, primarily by filtering unsafe data or rewriting it into safer forms.

arXiv AI
Sep 17

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

The paper introduces "cunning questions"—non‑safety prompts that contain misleading premises or subtle inconsistencies—to train large language models (LLMs) to scrutinize underlying intent and assumptions. Experiments show that incorporating these questions improves robustness against out‑of‑distribution jailbreak attacks and enhances subsequent safety fine‑tuning, achieving a new state‑of‑the‑art reduction in mean ASR from 17.40% to 15.05% across nine backbone–benchmark combinations. The authors argue that this training fosters vigilance, enabling models to prioritize safety judgments before engaging in harmful planning.

By Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi
arXiv Computation and Language
Aug 25

Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks

arXiv:2506.07356v3 Announce Type: replace Abstract: While Finetuning-as-a-Service (FaaS) enables customization of Large Language Models (LLMs) using user data, this service is vulnerable to safety de...

By Seokil Ham, Yubin Choi, Yujin Yang, Seungju Cho, Younghun Kim, Changick Kim
arXiv Machine Learning
Sep 22

Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs

The paper introduces MMSAFE, a multi-layer framework designed to identify safety-degrading data in multilingual large language models. It shows that safety signals are distributed across multiple layers and only partially shared across languages, unlike the single-layer assumption used in monolingual settings. Experiments demonstrate that MMSAFE reduces harmful-response rates by 60% compared to random filtering and outperforms the best single-layer baseline across various models, languages, and safety benchmarks.

By Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan