arXiv Machine Learning By Yonghui Yang, Yihui Wang, Junwei Li, Jilong Liu, Fengbin Zhu, Weibiao Huang, Le Wu, Richang Hong, Tat-Seng Chua

Controllable Value Alignment in Large Language Models through Neuron-Level Editing

Read the original on arXiv Machine Learning →

arXiv:2602. 07356v2 Announce Type: replace Abstract: Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making expands.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

Training-Free Task Vectors for LLM Behavioral Control

The paper introduces Training-Free Task Vectors (TFTVs), a method for computing task-vector-like directions in large language models without fine‑tuning. TFTVs map activation steering vectors to rank‑one weight‑space edits using only forward‑pass statistics, enabling arithmetic operations such as learning, forgetting, and composing edits. Experiments show that TFTVs consistently amplify, suppress, and combine target behaviors while preserving general knowledge, outperforming other editing and steering baselines.

By Gabriel J. Perin, Lucas Boscaini, Andr\'e Araujo, Nina S. T. Hirata
arXiv Machine Learning
5d ago

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

The paper introduces AIMES, a framework for adaptive multi-value activation steering in large language models. AIMES builds layer‑specific bipolar directions for moral‑foundation values and uses intermediate‑layer vocabulary readouts as online observers to guide a controller that adjusts intervention strengths at each decoding step. Experiments across instruction‑tuned model families show that AIMES achieves depth‑dependent advantages over fixed joint steering and prompt‑based steering, with smaller activation‑space interventions and comparable response quality.

By Payel Bhattacharjee, Ravi Tandon
arXiv Machine Learning
Jul 15

Inference-Time Machine Unlearning via Gated Activation Redirection

arXiv:2605. 12765v3 Announce Type: replace Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety.

By Vin\'icius Conte Turani, Ot\'avio Parraga, Jo\~ao Vitor Boer Abitante, Kristen K. Arguello, Joana Pasquali, Ramiro N. Barros, Flavio du Pin Calmon, Christian Mattjie, Rodrigo C. Barros, Lucas S. Kupssinsk\"u
arXiv Machine Learning
Jul 16

Value Drifts: Tracing Value Alignment During LLM Post-Training

arXiv:2510. 26707v2 Announce Type: replace-cross Abstract: As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems.

By Mehar Bhatia, Shravan Nayak, Gaurav Kamath, Marius Mosbach, Karolina Sta\'nczak, Vered Shwartz, Siva Reddy