arXiv Machine Learning By Payel Bhattacharjee, Ravi Tandon

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

Read the original on arXiv Machine Learning →

The paper introduces AIMES, a framework for adaptive multi-value activation steering in large language models. AIMES builds layer‑specific bipolar directions for moral‑foundation values and uses intermediate‑layer vocabulary readouts as online observers to guide a controller that adjusts intervention strengths at each decoding step. Experiments across instruction‑tuned model families show that AIMES achieves depth‑dependent advantages over fixed joint steering and prompt‑based steering, with smaller activation‑space interventions and comparable response quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

arXiv:2609.06289v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...

By Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi