arXiv:2609.14151v1 Announce Type: cross
Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model b...
By Prajjwal Bhattarai, Tuka Alhanai
arXiv:2606. 08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors.
By Tuc Nguyen, Thai Le
arXiv:2602. 00945v2 Announce Type: replace-cross Abstract: LLMs are multilingual by training, yet their lingua franca is often English, reflecting English language dominance in pretraining.
By Anusa Saha, Tanmay Joshi, Vinija Jain, Aman Chadha, Amitava Das
arXiv:2609.31122v1 Announce Type: new
Abstract: Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate...
By Irene Tallini, Lorenzo Basile, Valentino Maiorca, Francesco Locatello, Alberto Cazzaniga
arXiv:2605. 05983v2 Announce Type: replace Abstract: Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones.
By Yuntai Bao, Qinfeng Li, Xinyan Yu, Ge Su, Wenqi Zhang, Liu Yan, Haiqin Weng, Jianwei Yin, Xuhong Zhang
arXiv:2609.07037v1 Announce Type: new
Abstract: Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional...
By Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi, Hisashi Kashima
arXiv:2605.12890v2 Announce Type: replace-cross
Abstract: The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-writte...
By Luxu Liang, Xiang Li
MetaSteer is a new method for steering large language models that learns nonlinear, context-dependent interventions applied to attention projection matrices. Unlike traditional linear, context-independent techniques, MetaSteer adapts its effects based on the input, requiring no linear concept-geometry assumption. Trained once on a pooled preference corpus, it transfers zero‑shot to unseen concepts and out‑of‑distribution contexts, matching or surpassing strong task‑specific baselines on multiple benchmarks and model families.
By Mehdi Jafari, Hao Xue, Flora Salim
arXiv:2607. 07128v1 Announce Type: new Abstract: Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks.
By Maximilian S. Ernst (Max Planck School of Cognition, Center for Lifespan Psychology Max Planck Institute for Human Development, Machine Learning Group Technische Universit\"at Berlin), Lorenz Linhardt (Machine Learning Group Technische Universit\"at Berlin, Berlin Institute for the Foundations of Learning and Data), Aaron Peikert (Center for Lifespan Psychology Max Planck Institute for Human Development), Oliver Eberle (Machine Learning Group Technische Universit\"at Berlin, Berlin Institute for the Foundations of Learning and Data)
The paper introduces Minimally Invasive Steering Vector Optimization (MISVO), a method that adjusts a frozen language model’s final hidden states by adding vectors to steer outputs toward a test‑time reward while penalizing changes using the local KL geometry of the token distribution. MISVO derives an analytic Fisher term and a suffix score‑function term, showing that the suffix term is second‑order and that Fisher surrogates match the full KL gradient to first order. Experiments on preference and code‑generation tasks with 1B–14B parameter models demonstrate that MISVO achieves the highest mean reward in most settings while maintaining diversity and coherence comparable to Best‑of‑N.
By Taha Entesari, Jingyu Zhang, Daniel Khashabi, Mahyar Fazlyab
IDEA is a training‑free, input‑dependent steering method for large language models that matches activations to cluster‑specific directions aligned with a target concept. It clusters positive and negative activation supports per attention head, solves an optimal‑matching problem to create a pool of cluster‑conditional directions, and selects the best match for each input at inference time. This approach preserves the input’s original representation while improving the truth × info rate on TruthfulQA by an average of 9.9% (up to 23.5%) over input‑independent baselines.
By Zheng Wang, Muchen Li, Renjie Liao, Yan Leng
IDEA: training-free Input-Dependent stEEring via Activation cluster matching (IDEEA) is a method that steers large language models by injecting bias into selected activations at inference time, without requiring weight updates. Unlike existing training-free steering approaches that use a single, input-independent direction, IDEEA clusters positive and negative activation supports per attention head and solves an optimal-matching problem to create a set of cluster-conditional directions. At inference, IDEEA selects the direction that best matches the input’s activation, aligning the model toward a target concept while preserving the input’s original representation, and achieves a 9.9% average improvement in truth × info rate on TruthfulQA compared to the best input-independent baseline.