arXiv AI

Refusal Lives Downstream of Persona in Chat Models

arXiv:2606. 26161v1 Announce Type: new Abstract: Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been studied as separate mechanisms.

arXiv Computation and Language
Sep 16

There Is More to Refusal in Large Language Models than a Single Direction

The paper challenges the notion that refusal in large language models is governed by a single direction in activation space. It demonstrates that different refusal and non‑compliance categories map to distinct geometric directions, yet steering along any of these directions yields similar refusal–over‑refusal trade‑offs, acting as a shared one‑dimensional control knob. Using sparse autoencoders, the authors reveal a structured internal representation of refusal, comprising a reusable core of shared latents and style‑ or domain‑specific latents, and show that linear interventions collapse this structure into uniform behavioral control.

By Faaiz Joad, Majd Hawasly, Sabri Boughorbel, Nadir Durrani, Husrev Taha Sencar
arXiv AI
4d ago

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

The paper introduces Persona Dosing, a method that uses an activation‑steering coefficient to control the intensity of a language model’s persona traits. By conditioning a FLAS controller on trait descriptions and calibrating its flow time against measured trait expression, the approach can adjust trait intensity without requiring paired training data. Experiments on Llama‑3.1‑8B, Qwen3‑8B, and Gemma‑3‑4B show significant increases in core‑trait expression and low targeting errors across multiple traits.

By Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen, Jingyuan Zhang, Yuxuan Zhang, Xinjie Shen
arXiv Machine Learning
Sep 22

Causal Localization of the Refusal Direction in Audio Language Models

The study investigates where a large audio language model (LALM) derives its refusal responses to harmful spoken requests. By applying causal interventions at the audio-to-language-model interface and within the language-model residual layers, the authors find that the primary influence on the refusal margin comes from mid-to-late layers of the text language model rather than the audio front end. Ablations of the audio interface have minimal effect, while zeroing the encoder output still reduces the margin, indicating the audio pathway remains active but is not the main source of refusal decisions.

By Leonardo Haw-Yang Foo, Hung-yi Lee
arXiv Machine Learning
Aug 27

Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks

The paper investigates how refusal training shapes the internal geometry of language models, showing that activation updates from refusal-completion losses create a distinct low‑dimensional refusal subspace. In a case study on OLMo‑2‑0425‑1B‑Instruct, the authors link the brittleness of refusal directions to repetitive refusal prefixes and demonstrate that using diverse refusal starts can increase the stable rank of gradients, thereby hardening the model against vector‑ablation attacks. The work provides insights into the emergence of safety‑critical features and offers a potential strategy to strengthen refusal robustness.

By Andrey Labunets