arXiv Machine Learning By Yuxiao Li, Gjergji Kasneci

Safety Cost of Steering Vectors Is Separable and Reducible

Read the original on arXiv Machine Learning →

arXiv:2608. 08383v1 Announce Type: cross Abstract: Steering vectors are a lightweight tool for controlling LLM behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
4d ago

How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models

The paper introduces CAM-Steer, a Category‑Adaptive Multi‑category Safety Steering framework that estimates risk for each harm category by comparing hidden states to safe and unsafe prototypes. It then combines safety directions into a single steering vector and applies a rotation whose angle is set by the estimated risks, preserving the hidden‑state norm. Experiments on three LLM backbones and seven harm categories show that CAM‑Steer outperforms baselines in defense success rate, even when multiple harm categories co‑occur, with negligible inference overhead.

By Chenxi Wang, Ruiyang Huang, Li Huang, Yifan Wu