Stored in Optimizer State, Valued by Later Training: A Causal Account of Subliminal Trait Transfer
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates subliminal learning, where hidden traits from a teacher model are transferred to a student during distillation. It introduces trait‑direction drift as the underlying mechanism, showing that biased generation creates measurable preference gaps that accumulate into behavioral transfer during fine‑tuning. The authors propose probe‑space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction, significantly reducing hidden‑trait transfer while maintaining task performance.
arXiv:2606. 00995v1 Announce Type: new Abstract: Subliminal learning refers to a student language model acquiring a teacher's traits (e.
arXiv:2608.24593v1 Announce Type: new Abstract: Adaptive optimizers retain gradient history in moment variables, allowing a local change in loss weighting to alter later updates. We examine whether t...
arXiv:2605. 07284v2 Announce Type: replace Abstract: A late-layer change learned during post-training may work on the base model's earlier state, or it may depend on earlier computation learned with it.
The paper investigates how scaling the amount of model-generated, off‑task distillation data affects the recoverability of teacher‑induced traits in student models. In a controlled subliminal‑learning setup, teachers are prompted to express a target trait, producing restricted data such as number‑only completions. Students trained on larger independent datasets show a clearer manifestation of the teacher’s trait in a separate evaluation domain, with the effect being strongest when the trait is already favored or when alternative traits are present. The authors find this trend holds across model families, trait types, multi‑trait settings, and cross‑model transfer, and suggest that scaling should be coupled with trait‑aware curation and evaluation.
arXiv:2607. 11958v1 Announce Type: new Abstract: Under the free energy principle, a predictive system does not observe reality directly; it maintains a generative model of the world and experiences that model's best current hypothesis.