arXiv:2606. 19168v1 Announce Type: new Abstract: To achieve deeper safety alignment for large language models (LLMs), recent efforts have studied how to push safety interventions earlier into the pretraining stage, primarily by filtering unsafe data or rewriting it into safer forms.
By Jinhan Li, Kexian Tang, Yihan Xu, Zhuorui Ye, Kaifeng Lyu
arXiv:2511. 00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks.
By Mina Taraghi, Yann Pequignot, Amin Nikanjam, Mohamed Amine Merzouk, Foutse Khomh
Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them.
arXiv:2606. 02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax.
By Hao Li, Jingkun An, Zijun Song, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha
SHARD is a self‑reframing distillation technique designed to enhance the safe‑helpfulness of large language models. It rewrites sensitive prompts to reveal benign intent, reframes the model’s original responses into safer, more helpful versions, and then fine‑tunes the model on these self‑reframed outputs. Experiments on DNA and the English subset of LINGUASAFE show that SHARD improves helpfulness across various model families while maintaining safety, and performs competitively with distillation from larger teacher models.
By Viswonathan Manoranjan, Amogh Gupta, Anvesh Rao Vijjini, Thomas Hofweber, Snigdha Chaturvedi
EOPSA (Efficient On-Policy Self-Distilled Safety Alignment) addresses inefficiencies in On-Policy Self-Distillation (OPSD) for safety alignment by focusing training on safety-critical tokens. It introduces Adaptive Rollout Scheduling, which limits generation length based on a Teacher Rescue Rate metric, and Selective Distillation, which filters out safety-neutral tokens to concentrate gradient updates on safety-pivotal transitions. Experiments on models up to 32B parameters show that EOPSA reduces rollout computation by about 50% and backpropagates through only roughly 2% of tokens, outperforming full-token distillation baselines in safety compliance and reasoning retention.