A Gravitational Interpretation of Fine-Tuning Reversion
arXiv:2606. 28525v1 Announce Type: cross Abstract: Fine-tuning on harmless data can partially undo behaviors acquired earlier in training.
The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.
arXiv:2606. 28525v1 Announce Type: cross Abstract: Fine-tuning on harmless data can partially undo behaviors acquired earlier in training.
arXiv:2605. 12765v3 Announce Type: replace Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety.
arXiv:2606. 19831v1 Announce Type: cross Abstract: Aligned language models gate behaviors such as refusal and language routing through sparse feed forward neurons, yet no theory predicts when a single neuron intervention controls a behavior coherently rather than collapsing the output.
arXiv:2608. 06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines.
arXiv:2605. 12705v2 Announce Type: replace Abstract: How can we train models whose post-trained capabilities survive subsequent fine-tuning?
arXiv:2608.29458v1 Announce Type: cross Abstract: Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluatio...
arXiv:2606. 08365v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can behave inconsistently across contexts and perturb unrelated features.
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
arXiv:2602. 06941v2 Announce Type: replace-cross Abstract: Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.
arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.
arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).
arXiv:2607. 25270v1 Announce Type: cross Abstract: Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail.