arXiv AI

SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives

arXiv:2509. 13450v3 Announce Type: replace Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets.

arXiv AI
Sep 10

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

arXiv:2609.06289v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...

By Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi
arXiv Machine Learning
4d ago

How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models

The paper introduces CAM-Steer, a Category‑Adaptive Multi‑category Safety Steering framework that estimates risk for each harm category by comparing hidden states to safe and unsafe prototypes. It then combines safety directions into a single steering vector and applies a rotation whose angle is set by the estimated risks, preserving the hidden‑state norm. Experiments on three LLM backbones and seven harm categories show that CAM‑Steer outperforms baselines in defense success rate, even when multiple harm categories co‑occur, with negligible inference overhead.

By Chenxi Wang, Ruiyang Huang, Li Huang, Yifan Wu
arXiv AI
Jul 3

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

arXiv:2510. 04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactions in socially interactive settings.

By Amin Banayeeanzade, Ala N. Tak, Fatemeh Bahrani, Anahita Bolourani, Leonardo Blas, Emilio Ferrara, Jonathan Gratch, Sai Praneeth Karimireddy