arXiv:2609.06289v1 Announce Type: cross
Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...
By Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi
The paper investigates how large language models (LLMs) share a common Fisher‑Rao geometry in their next‑token probability distributions, revealing that behaviour largely determines this geometry while activation geometry depends on coordinate choices. Across transformer, state‑space, and recurrent architectures, output geometries align more closely than activation geometries, and this shared structure facilitates semantic‑category transfer and improves agreement with human word choices as models scale and train. The study further demonstrates that geometry can guide minimum‑disturbance interventions, enabling reusable control that preserves behaviour better than Euclidean methods and enhances steering, editing, attribution, dictionary learning, and fine‑tuning.
By Dario Picozzi
GeoSteer introduces a geometry-aware, optimization-based approach to norm-preserving activation steering in large language models. By formulating steering as a Riemannian optimization problem, it updates activations through a sequence of small geodesic steps guided by a learned nonlinear objective, avoiding fixed steering directions. Experiments on TruthfulQA, RealToxicityPrompts, and UltraFeedback show that GeoSteer consistently outperforms existing activation steering baselines, offering smoother, more stable, and more consistent steering behavior.
By Xuan Cuong Ngo, Hao Vo, Ngan Le
arXiv:2607. 05790v1 Announce Type: new Abstract: Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily.
By Yuqi Chen, Vincent Siu, Yang Liu, Dawn Song, Chenguang Wang
arXiv:2607. 25270v1 Announce Type: cross Abstract: Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail.
By Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang, Lei Hou, Juanzi Li, Liangming Pan
arXiv:2608. 02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.
By Max Torop, Aria Masoomi, Jennifer Dy
arXiv:2609.14151v1 Announce Type: cross
Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model b...
By Prajjwal Bhattarai, Tuka Alhanai
arXiv:2606. 08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors.
By Tuc Nguyen, Thai Le
arXiv:2608. 11227v1 Announce Type: new Abstract: Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining.
By Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang, Jun Sun
arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).
By Qi Cao, Jian Lou, Meiting Liu, Wenjie Feng, Dan Li, See-Kiong Ng, Anh Tuan Luu
arXiv:2602. 02712v2 Announce Type: replace Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations.
By Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau
arXiv:2606. 15092v1 Announce Type: new Abstract: Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs).
By Minh-Hieu Pham, Bach Do, Laziz Abdullaev, Tan Minh Nguyen, Khoat Than