arXiv Machine Learning

From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models

The paper investigates how steering a language model along a trait vector—specifically for refusal and sycophancy—affects its internal computation. By splitting the model’s computation into a reconstruction circuit before the trait vector and a transmission circuit after it, the authors find that for refusal both circuits are compact and faithful, with the transmission circuit alone largely restoring the lost refusal signal. For sycophancy, the transmission circuit remains compact but the reconstruction circuit is broader and only partially faithful, indicating that the circuits used for steering are not necessarily the same as those that generate the behavior.

arXiv Machine Learning
Aug 27

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.

By Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron
arXiv AI
Sep 2

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

The paper investigates how steering vectors influence large language models (LLMs) by conducting a mechanistic case study on refusal behavior. Using a multi-token activation patching framework, the authors find that steering methods primarily target the OV circuit of the attention mechanism, largely ignoring the QK circuit, and that these circuits are functionally interchangeable across different steering approaches. The study also shows that steering vectors can be sparsified by 85–96% with minimal performance loss and that key dimensions are consistently identified across methods.

By Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha
arXiv Machine Learning
Sep 23

Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages

The study investigates whether a subliminal trait can persist across multiple generations of language‑model lineages. Three copies of Qwen2.5‑7B‑Instruct were trained for ten iterations, and the trait’s expression was measured via a keyword screen and an activation probe. Results show the trait remains detectable in all generations, though its behavioral expression diminishes, and it can exist internally without being overtly expressed when the system prompt is removed.

By Ryan Vo, Duc-Vu Nguyen, Matt Kretchmar, Ngan Luu-Thuy Nguyen
arXiv Computation and Language
Aug 28

Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery

The paper investigates whether structural differences in circuits discovered by circuit discovery methods reflect distinct mechanisms. By varying input-token frequency while keeping the task constant, the authors find that although circuits appear specialized by frequency structurally, functional and representational analyses reveal no reliable differences, a phenomenon they call phantom specialization. Across multiple models and tasks, structurally distinct circuits implement the same computation, with core shared subgraphs recovering most of the performance and interchangeable internal representations confirmed by causal interventions.

By Alireza Bayat Makou, Jingcheng Niu, Subhabrata Dutta, Iryna Gurevych
arXiv Computation and Language
Sep 4

How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

The study investigates the consistency and specificity of language model circuits across six tasks and five models, focusing on component-level (attention heads and MLP blocks) and neuron-level circuits. Component-level circuits are highly consistent and causally important but lack task specificity, as ablating a circuit for one task similarly harms performance on other tasks. Neuron-level circuits show higher task specificity but lower consistency, with overlap mainly between closely related tasks. The analysis of Llama‑3.2‑3B reveals that shared components are predominantly MLP blocks, while attention heads act as generic attention‑sink heads.

By Michael Li, Nishant Subramani
arXiv AI
6d ago

FLIP: Final Layer Inference-Time Probing for Vision-Language Models

FLIP is a final‑layer inference‑time probe designed to test whether a logit‑facing intervention site in an open‑weight vision‑language model (VLM) supports structured, task‑linked computation rather than generic perturbation. The probe applies elementwise flooring to the final normalized hidden state before logit computation, leaving other model components unchanged. By sweeping intervention strength on a controlled detection/counting task, FLIP identifies three regimes—negligible change, a bounded interior regime with improved detection recall and reduced counting error, and over‑suppression—while a four‑criterion protocol ensures the observed effects are mechanistically interpretable.

By Drandreb Earl O. Juanico, Rowel O. Atienza