arXiv AI By Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja, Saiyam Jain, Vaishnavi Shivkumar, Ponnurangam Kumaraguru

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

Read the original on arXiv AI →

arXiv:2607. 19806v1 Announce Type: cross Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors may induce over-refusal on benign prompts.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.