arXiv Machine Learning By Chenxi Wang, Ruiyang Huang, Li Huang, Yifan Wu

How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models

Read the original on arXiv Machine Learning →

The paper introduces CAM-Steer, a Category‑Adaptive Multi‑category Safety Steering framework that estimates risk for each harm category by comparing hidden states to safe and unsafe prototypes. It then combines safety directions into a single steering vector and applies a rotation whose angle is set by the estimated risks, preserving the hidden‑state norm. Experiments on three LLM backbones and seven harm categories show that CAM‑Steer outperforms baselines in defense success rate, even when multiple harm categories co‑occur, with negligible inference overhead.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 18

The Role of Fine-grained Harm Signals in LLM Safety

The study investigates how category‑specific harmfulness signals, isolated by removing the shared general harmfulness component, influence large language model (LLM) safety. Using activation steering across 11 risk categories in three instruction‑tuned LLMs, the authors find that the presence of harmfulness in these category residuals varies by category and that the pattern of inducing refusal is even more model‑dependent. Additionally, category residuals were shown to enhance the models’ downstream alignment with the shared general harmfulness representation, indicating that fine‑grained signals play a role beyond the general component.

By Soyeon Park (KAIST), Seogyeong Jeong (KAIST), Sunwoo Kim (KAIST), Alice Oh (KAIST)