arXiv AI By Alex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab, Murat Cubuktepe, Mike Vaiana, Diogo de Lucena, Judd Rosenblatt, Michael S. A. Graziano

Endogenous Resistance to Activation Steering in Language Models

Read the original on arXiv AI →

arXiv:2602. 06941v2 Announce Type: replace-cross Abstract: Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.