GeoSteer introduces a geometry-aware, optimization-based approach to norm-preserving activation steering in large language models. By formulating steering as a Riemannian optimization problem, it updates activations through a sequence of small geodesic steps guided by a learned nonlinear objective, avoiding fixed steering directions. Experiments on TruthfulQA, RealToxicityPrompts, and UltraFeedback show that GeoSteer consistently outperforms existing activation steering baselines, offering smoother, more stable, and more consistent steering behavior.
By Xuan Cuong Ngo, Hao Vo, Ngan Le
FishBack introduces a pullback Fisher geometry approach for activation steering in transformers, challenging the common Euclidean assumption of intermediate activation spaces. By deriving a closed‑form steering direction based on the Fisher information metric of the softmax layer, the method achieves target concept changes with minimal off‑target distortion, especially in early and middle layers. Experiments on GPT‑2 Small, Llama‑3‑8B, and Qwen3‑8B demonstrate significant reductions in off‑target KL divergence compared to existing steering baselines.
By Sihan Wang, Jiayi Zhao, Qingyan Cao, Hongbo Yao, Lin Shu
arXiv:2606. 08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors.
By Tuc Nguyen, Thai Le
IDEA is a training‑free, input‑dependent steering method for large language models that matches activations to cluster‑specific directions aligned with a target concept. It clusters positive and negative activation supports per attention head, solves an optimal‑matching problem to create a pool of cluster‑conditional directions, and selects the best match for each input at inference time. This approach preserves the input’s original representation while improving the truth × info rate on TruthfulQA by an average of 9.9% (up to 23.5%) over input‑independent baselines.
By Zheng Wang, Muchen Li, Renjie Liao, Yan Leng
arXiv:2607. 19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection.
By Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan
arXiv:2606. 18844v1 Announce Type: new Abstract: Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution.
By Zhilin Huang, Hang Gao, Ziqiang Dong, Yuan Chen, Yifeng Luo, Chujun Qin, Jingyi Wang, Yang Yang, Guanjun Jiang