arXiv AI

A Gravitational Interpretation of Fine-Tuning Reversion

arXiv:2606. 28525v1 Announce Type: cross Abstract: Fine-tuning on harmless data can partially undo behaviors acquired earlier in training.

arXiv Machine Learning
Aug 27

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.

By Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron
arXiv AI
Jul 21

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv:2607. 18114v1 Announce Type: cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer.

By Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch\"olkopf, Zhijing Jin
Hugging Face Trending Papers
Jul 20

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model.

arXiv AI
Sep 24

Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?

The paper investigates what a steerable valence direction in large language models (LLMs) actually represents, focusing on a good‑bad outcome direction in a maze task. By using controlled interventions that separate the realized outcome from the informational history that led to it, the authors find that directions trained on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer performance drops sharply, indicating that the post‑event readout remains strongly conditioned on the earlier announcement. In a matched maze‑reinforcement‑learning run, the post‑RL direction becomes more predictive of reference‑MDP return and the policy depends more on it, yet the history dependence persists. These findings support a functional, value‑related interpretation of the direction but argue against identifying it with a history‑invariant scalar valence state.

By Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng