arXiv:2502. 12446v3 Announce Type: replace-cross Abstract: Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.
By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
arXiv:2602.01654v2 Announce Type: replace
Abstract: Steering vectors (SVs) offer a lightweight way to control large language models (LLMs) at inference time by shifting hidden activations, providing...
By Jiaqian Li, Yanshu Li, Kuan-Hao Huang
The paper introduces Training-Free Task Vectors (TFTVs), a method for computing task-vector-like directions in large language models without fine‑tuning. TFTVs map activation steering vectors to rank‑one weight‑space edits using only forward‑pass statistics, enabling arithmetic operations such as learning, forgetting, and composing edits. Experiments show that TFTVs consistently amplify, suppress, and combine target behaviors while preserving general knowledge, outperforming other editing and steering baselines.
By Gabriel J. Perin, Lucas Boscaini, Andr\'e Araujo, Nina S. T. Hirata
arXiv:2609.06289v1 Announce Type: cross
Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...
By Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi
arXiv:2509. 25533v2 Announce Type: replace-cross Abstract: As Vision Language Models (VLMs) are deployed across safety-critical applications, understanding and controlling their behavioral patterns has become increasingly important.
By Ravikumar Balakrishnan, Mansi Phute
arXiv:2607. 19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection.
By Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan
The paper introduces AIMES, a framework for adaptive multi-value activation steering in large language models. AIMES builds layer‑specific bipolar directions for moral‑foundation values and uses intermediate‑layer vocabulary readouts as online observers to guide a controller that adjusts intervention strengths at each decoding step. Experiments across instruction‑tuned model families show that AIMES achieves depth‑dependent advantages over fixed joint steering and prompt‑based steering, with smaller activation‑space interventions and comparable response quality.
By Payel Bhattacharjee, Ravi Tandon
IDEA is a training‑free, input‑dependent steering method for large language models that matches activations to cluster‑specific directions aligned with a target concept. It clusters positive and negative activation supports per attention head, solves an optimal‑matching problem to create a pool of cluster‑conditional directions, and selects the best match for each input at inference time. This approach preserves the input’s original representation while improving the truth × info rate on TruthfulQA by an average of 9.9% (up to 23.5%) over input‑independent baselines.
By Zheng Wang, Muchen Li, Renjie Liao, Yan Leng
arXiv:2604. 03532v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult.
By Sing Hieng Wong, Hassan Sajjad, A. B. Siddique
arXiv:2605. 05983v2 Announce Type: replace Abstract: Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones.
By Yuntai Bao, Qinfeng Li, Xinyan Yu, Ge Su, Wenqi Zhang, Liu Yan, Haiqin Weng, Jianwei Yin, Xuhong Zhang
arXiv:2607. 03522v1 Announce Type: new Abstract: Fine-tuning a single low-rank adapter on many domains at once is multi-task learning: the domains must be co-learned, and how they share the adapter decides whether they help or hurt one another.
By Wei Zhang, Lin Tang, Ming Zhao, Yuxuan Wang
arXiv:2601. 09361v4 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a key paradigm for improving large-scale reasoning models.
By Jiaying Zhang, Lei Shi, Jiguo Li, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He