arXiv Machine Learning By Bochen Lyu, Yiyang Jia, Xiaohao Cai, Zhanxing Zhu

When Autoregressive Consistency Hurts Safety Alignment

Read the original on arXiv Machine Learning →

arXiv:2606. 04168v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near the first few output tokens.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.