arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.
By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou
Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Finding the regime and boundaries of successful steering typically requires expensive grid searches and post-hoc evaluation of full autoregressive rollouts.
arXiv:2602.01654v2 Announce Type: replace
Abstract: Steering vectors (SVs) offer a lightweight way to control large language models (LLMs) at inference time by shifting hidden activations, providing...
By Jiaqian Li, Yanshu Li, Kuan-Hao Huang
The paper introduces Minimally Invasive Steering Vector Optimization (MISVO), a method that adjusts a frozen language model’s final hidden states by adding vectors to steer outputs toward a test‑time reward while penalizing changes using the local KL geometry of the token distribution. MISVO derives an analytic Fisher term and a suffix score‑function term, showing that the suffix term is second‑order and that Fisher surrogates match the full KL gradient to first order. Experiments on preference and code‑generation tasks with 1B–14B parameter models demonstrate that MISVO achieves the highest mean reward in most settings while maintaining diversity and coherence comparable to Best‑of‑N.
By Taha Entesari, Jingyu Zhang, Daniel Khashabi, Mahyar Fazlyab
arXiv:2502. 12446v3 Announce Type: replace-cross Abstract: Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.
By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
arXiv:2602. 02712v2 Announce Type: replace Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations.
By Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau
MetaSteer is a new method for steering large language models that learns nonlinear, context-dependent interventions applied to attention projection matrices. Unlike traditional linear, context-independent techniques, MetaSteer adapts its effects based on the input, requiring no linear concept-geometry assumption. Trained once on a pooled preference corpus, it transfers zero‑shot to unseen concepts and out‑of‑distribution contexts, matching or surpassing strong task‑specific baselines on multiple benchmarks and model families.
By Mehdi Jafari, Hao Xue, Flora Salim
arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.
By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
arXiv:2606. 07696v1 Announce Type: cross Abstract: Activation steering has become a popular training-free method to control LLM behavior by injecting precomputed direction vectors into the model's residual stream at inference time.
By Kien Le, Thai Le
arXiv:2606. 15092v1 Announce Type: new Abstract: Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs).
By Minh-Hieu Pham, Bach Do, Laziz Abdullaev, Tan Minh Nguyen, Khoat Than
arXiv:2606.12234v2 Announce Type: replace
Abstract: Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the invol...
By Iuri Macocco, Pau Rodr\'iguez, Arno Blaas, Luca Zappella, Marco Baroni, Xavier Suau
arXiv:2509. 03647v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models.
By Dani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas, Simon Fu, Narmeen Oozeer