The paper introduces AIMES, a framework for adaptive multi-value activation steering in large language models. AIMES builds layer‑specific bipolar directions for moral‑foundation values and uses intermediate‑layer vocabulary readouts as online observers to guide a controller that adjusts intervention strengths at each decoding step. Experiments across instruction‑tuned model families show that AIMES achieves depth‑dependent advantages over fixed joint steering and prompt‑based steering, with smaller activation‑space interventions and comparable response quality.
By Payel Bhattacharjee, Ravi Tandon
arXiv:2502. 12446v3 Announce Type: replace-cross Abstract: Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.
By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
arXiv:2602. 03160v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles.
By Woojin Kim, Sieun Hyeon, Jusang Oh, Jaeyoung Do
arXiv:2602.01654v2 Announce Type: replace
Abstract: Steering vectors (SVs) offer a lightweight way to control large language models (LLMs) at inference time by shifting hidden activations, providing...
By Jiaqian Li, Yanshu Li, Kuan-Hao Huang
arXiv:2608. 02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.
By Max Torop, Aria Masoomi, Jennifer Dy
The paper investigates whether large language models (LLMs) possess intrinsic value systems and how to quantify and align them. By projecting responses from 106 LLMs and 95,000 human survey profiles into a shared sociological space, the authors confirm that LLMs do have values, though these values form a concentrated, idealized core rather than mirroring human diversity. They introduce the Prior-Environment-Cognition (PEC) framework to mathematically define value expression and propose an adaptive Alignment Prescription that identifies minimal interventions—ranging from prompts to targeted parameter updates—to steer LLM values efficiently without harming general performance.
By Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu
arXiv:2602. 07356v2 Announce Type: replace Abstract: Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making expands.
By Yonghui Yang, Yihui Wang, Junwei Li, Jilong Liu, Fengbin Zhu, Weibiao Huang, Le Wu, Richang Hong, Tat-Seng Chua
arXiv:2609.39701v1 Announce Type: new
Abstract: Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Convention...
By Jiale Dai, Hongcan Deng, Liuxian Ma, Xiaoke Niu, Guojie Song
arXiv:2607. 18259v1 Announce Type: new Abstract: Steering vectors (SVs), an inference-time intervention technique for large language models (LLMs), guide the generation process by adding a concept-specific direction vector to intermediate activations during inference.
By Brian Becker, Rui Chu, Yingjie Lao
arXiv:2510. 01167v2 Announce Type: replace-cross Abstract: Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a single objective.
By Yiran Shen, Yu Xia, Jonathan Chang, Prithviraj Ammanabrolu
arXiv:2609.07037v1 Announce Type: new
Abstract: Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional...
By Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi, Hisashi Kashima
arXiv:2609.06289v1 Announce Type: cross
Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...
By Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi