arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.
By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
arXiv:2609.07037v1 Announce Type: new
Abstract: Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional...
By Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi, Hisashi Kashima
arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.
By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou
arXiv:2605.12890v2 Announce Type: replace-cross
Abstract: The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-writte...
By Luxu Liang, Xiang Li
Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Finding the regime and boundaries of successful steering typically requires expensive grid searches and post-hoc evaluation of full autoregressive rollouts.
arXiv:2607. 25270v1 Announce Type: cross Abstract: Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail.
By Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang, Lei Hou, Juanzi Li, Liangming Pan
arXiv:2608. 05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful.
By Bohan Jiang, Dawei Li, Yasin Silva, Huan Liu
arXiv:2602. 02712v2 Announce Type: replace Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations.
By Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau
arXiv:2603.13359v2 Announce Type: replace
Abstract: Language models are commonly fine-tuned for safety alignment to refuse harmful prompts. One approach fine-tunes them to generate categorical refusa...
By Rishab Alagharu, Ishneet Sukhvinder Singh, Shaibi Shamsudeen, Zhen Wu, Ashwinee Panda
arXiv:2609.24821v1 Announce Type: new
Abstract: The Linear Representation Hypothesis associates high-level concepts with directions in language models, but it remains unclear how these concept-relate...
By Manjiang Yu, Hongji Li, Zihan Wang, Junwei Chen, Xue Li, Priyanka Singh, Yang Cao, Lijie Hu
The paper investigates what aspects of language model behavior are controlled by activation steering. By introducing Cross‑Encoding Steering Evaluation, the authors show that steering effects often follow the extraction index of answer identifiers rather than the semantic content of the answers, especially at deeper layers. They also find that a low‑rank output‑sensitive component captures most of this effect, and that different datasets (NormBank, MNLI, SC101) exhibit varying preferences for extraction‑index versus semantic‑label following.
By Zhiwei Gao, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki
Activation steering is often evaluated under the answer encoding used to construct the direction. A reported gain may reflect the intended judgment or compatibility with answer identifiers seen during...