arXiv Machine Learning By Zhichen Dong, Zhixuan Liu, Yuyu Fan, Xiangtian Li, Shuyang Zhang, Chao Yang

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

Read the original on arXiv Machine Learning →

The paper investigates how scaling the amount of model-generated, off‑task distillation data affects the recoverability of teacher‑induced traits in student models. In a controlled subliminal‑learning setup, teachers are prompted to express a target trait, producing restricted data such as number‑only completions. Students trained on larger independent datasets show a clearer manifestation of the teacher’s trait in a separate evaluation domain, with the effect being strongest when the trait is already favored or when alternative traits are present. The authors find this trend holds across model families, trait types, multi‑trait settings, and cross‑model transfer, and suggest that scaling should be coupled with trait‑aware curation and evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
2d ago

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

The paper investigates subliminal learning, where hidden traits from a teacher model are transferred to a student during distillation. It introduces trait‑direction drift as the underlying mechanism, showing that biased generation creates measurable preference gaps that accumulate into behavioral transfer during fine‑tuning. The authors propose probe‑space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction, significantly reducing hidden‑trait transfer while maintaining task performance.

By Zhixuan Liu, Zhichen Dong, Yuyu Fan, Xiangtian Li, Chao Yang
arXiv Machine Learning
Jul 14

Reference-Based Distillation Detection in LLMs

arXiv:2607. 09692v1 Announce Type: new Abstract: Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations.

By Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, Sewon Min