arXiv AI By Hao Li, Jingkun An, Zijun Song, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

Read the original on arXiv AI →

arXiv:2606. 02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.