The paper investigates how scaling the amount of model-generated, off‑task distillation data affects the recoverability of teacher‑induced traits in student models. In a controlled subliminal‑learning setup, teachers are prompted to express a target trait, producing restricted data such as number‑only completions. Students trained on larger independent datasets show a clearer manifestation of the teacher’s trait in a separate evaluation domain, with the effect being strongest when the trait is already favored or when alternative traits are present. The authors find this trend holds across model families, trait types, multi‑trait settings, and cross‑model transfer, and suggest that scaling should be coupled with trait‑aware curation and evaluation.
By Zhichen Dong, Zhixuan Liu, Yuyu Fan, Xiangtian Li, Shuyang Zhang, Chao Yang
arXiv:2606. 00995v1 Announce Type: new Abstract: Subliminal learning refers to a student language model acquiring a teacher's traits (e.
By Camila Blank, Agam Bhatia, Senthooran Rajamanoharan, Arthur Conmy, Neel Nanda
arXiv:2608.20442v1 Announce Type: new
Abstract: Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically...
By Qinyang Xu
arXiv:2608. 05734v1 Announce Type: new Abstract: Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models.
By Ethan Hadley, Eren Gultepe
Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a teacher model to a student by distilling from seemingly unrelated or random synthetic data from the teacher.
arXiv:2608. 03632v1 Announce Type: new Abstract: On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals.
By Yinuo Jiang, Yongjie Ye, Zhou Tao, Xiang Zhuang, Qiang Zhang, Huajun Chen, Tiankai Li
arXiv:2606. 12171v1 Announce Type: cross Abstract: Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in probability distributions, and mixup enforces them through convex combinations of inputs.
By Jos\'e Medina, Paul Honeine, Abdelaziz Bensrhair, Amnir Hadachi
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable.
arXiv:2608. 15787v1 Announce Type: cross Abstract: Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts.
By Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim
arXiv:2607. 04432v1 Announce Type: cross Abstract: A student model trained on pure uniform noise can still inherit its teacher's digit-classification ability, provided the two share initialization.
By Kargi Chauhan, Aditya Shah
arXiv:2606. 28525v1 Announce Type: cross Abstract: Fine-tuning on harmless data can partially undo behaviors acquired earlier in training.
By Samuele Poppi, Nils Lukas
The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.
By Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron