arXiv Machine Learning By Oscar Mir\'o L\'opez-Feliu, Maya Ozbayoglu

From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models

Read the original on arXiv Machine Learning →

The paper investigates how steering a language model along a trait vector—specifically for refusal and sycophancy—affects its internal computation. By splitting the model’s computation into a reconstruction circuit before the trait vector and a transmission circuit after it, the authors find that for refusal both circuits are compact and faithful, with the transmission circuit alone largely restoring the lost refusal signal. For sycophancy, the transmission circuit remains compact but the reconstruction circuit is broader and only partially faithful, indicating that the circuits used for steering are not necessarily the same as those that generate the behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.

By Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron
arXiv AI
Sep 2

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

The paper investigates how steering vectors influence large language models (LLMs) by conducting a mechanistic case study on refusal behavior. Using a multi-token activation patching framework, the authors find that steering methods primarily target the OV circuit of the attention mechanism, largely ignoring the QK circuit, and that these circuits are functionally interchangeable across different steering approaches. The study also shows that steering vectors can be sparsified by 85–96% with minimal performance loss and that key dimensions are consistently identified across methods.

By Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha
arXiv Machine Learning
Sep 23

Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages

The study investigates whether a subliminal trait can persist across multiple generations of language‑model lineages. Three copies of Qwen2.5‑7B‑Instruct were trained for ten iterations, and the trait’s expression was measured via a keyword screen and an activation probe. Results show the trait remains detectable in all generations, though its behavioral expression diminishes, and it can exist internally without being overtly expressed when the system prompt is removed.

By Ryan Vo, Duc-Vu Nguyen, Matt Kretchmar, Ngan Luu-Thuy Nguyen