arXiv AI By Viola Zhong, Qirui Li

Refusal Lives Downstream of Persona in Chat Models

Read the original on arXiv AI →

arXiv:2606. 26161v1 Announce Type: new Abstract: Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been studied as separate mechanisms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 16

There Is More to Refusal in Large Language Models than a Single Direction

The paper challenges the notion that refusal in large language models is governed by a single direction in activation space. It demonstrates that different refusal and non‑compliance categories map to distinct geometric directions, yet steering along any of these directions yields similar refusal–over‑refusal trade‑offs, acting as a shared one‑dimensional control knob. Using sparse autoencoders, the authors reveal a structured internal representation of refusal, comprising a reusable core of shared latents and style‑ or domain‑specific latents, and show that linear interventions collapse this structure into uniform behavioral control.

By Faaiz Joad, Majd Hawasly, Sabri Boughorbel, Nadir Durrani, Husrev Taha Sencar
arXiv AI
4d ago

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

The paper introduces Persona Dosing, a method that uses an activation‑steering coefficient to control the intensity of a language model’s persona traits. By conditioning a FLAS controller on trait descriptions and calibrating its flow time against measured trait expression, the approach can adjust trait intensity without requiring paired training data. Experiments on Llama‑3.1‑8B, Qwen3‑8B, and Gemma‑3‑4B show significant increases in core‑trait expression and low targeting errors across multiple traits.

By Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen, Jingyuan Zhang, Yuxuan Zhang, Xinjie Shen