arXiv AI By Srikanth Malla, Chiho Choi, Joon Hee Choi

Steering Interference Reflects the Model's Defaults, Not the Behavior Directions

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Sep 22

From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models

The paper investigates how steering a language model along a trait vector—specifically for refusal and sycophancy—affects its internal computation. By splitting the model’s computation into a reconstruction circuit before the trait vector and a transmission circuit after it, the authors find that for refusal both circuits are compact and faithful, with the transmission circuit alone largely restoring the lost refusal signal. For sycophancy, the transmission circuit remains compact but the reconstruction circuit is broader and only partially faithful, indicating that the circuits used for steering are not necessarily the same as those that generate the behavior.

By Oscar Mir\'o L\'opez-Feliu, Maya Ozbayoglu