arXiv AI By Niklas Herbster, Martin Zborowski, Alberto Tosato, Gauthier Gidel, Tommaso Tosato

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

Read the original on arXiv AI →

arXiv:2604. 08169v2 Announce Type: replace Abstract: Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.