arXiv AI By Qi Cao, Jian Lou, Meiting Liu, Wenjie Feng, Dan Li, See-Kiong Ng, Anh Tuan Luu

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

Read the original on arXiv AI →

arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.