arXiv AI By Hao Li, Jingkun An, Zijun Song, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

Read the original on arXiv AI →

arXiv:2606. 02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

The paper introduces Suan, a new preference optimization algorithm designed to improve safety alignment in large language models. Suan operates directly at the gradient level, avoiding traditional variational derivations, which yields more interpretable and robust training dynamics. Experiments show that Suan outperforms existing methods, achieving superior safety alignment while maintaining response utility.

By Oleksandr Cherednichenko, Roman Klypa