arXiv AI

A Geometric Account of Activation Steering through Angle-Norm Decomposition

arXiv:2606. 06735v1 Announce Type: new Abstract: Linear activation steering has gained popularity as a simple and empirically effective way to control language model behavior.

arXiv Machine Learning
Jul 9

Towards Understanding Steering Strength

arXiv:2602. 02712v2 Announce Type: replace Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations.

By Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau