arXiv:2609.06289v1 Announce Type: cross
Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...
By Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi
The paper investigates how large language models (LLMs) share a common Fisher‑Rao geometry in their next‑token probability distributions, revealing that behaviour largely determines this geometry while activation geometry depends on coordinate choices. Across transformer, state‑space, and recurrent architectures, output geometries align more closely than activation geometries, and this shared structure facilitates semantic‑category transfer and improves agreement with human word choices as models scale and train. The study further demonstrates that geometry can guide minimum‑disturbance interventions, enabling reusable control that preserves behaviour better than Euclidean methods and enhances steering, editing, attribution, dictionary learning, and fine‑tuning.
By Dario Picozzi
GeoSteer introduces a geometry-aware, optimization-based approach to norm-preserving activation steering in large language models. By formulating steering as a Riemannian optimization problem, it updates activations through a sequence of small geodesic steps guided by a learned nonlinear objective, avoiding fixed steering directions. Experiments on TruthfulQA, RealToxicityPrompts, and UltraFeedback show that GeoSteer consistently outperforms existing activation steering baselines, offering smoother, more stable, and more consistent steering behavior.
By Xuan Cuong Ngo, Hao Vo, Ngan Le
arXiv:2607. 05790v1 Announce Type: new Abstract: Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily.
By Yuqi Chen, Vincent Siu, Yang Liu, Dawn Song, Chenguang Wang
arXiv:2607. 25270v1 Announce Type: cross Abstract: Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail.
By Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang, Lei Hou, Juanzi Li, Liangming Pan
arXiv:2608. 02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.
By Max Torop, Aria Masoomi, Jennifer Dy