arXiv Machine Learning By Bochen Lyu, Yiyang Jia, Xiaohao Cai, Zhanxing Zhu

When Autoregressive Consistency Hurts Safety Alignment

Read the original on arXiv Machine Learning →

arXiv:2606. 04168v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near the first few output tokens.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.