arXiv AI

Covert Trait Propagation Is Representation Alignment: Mechanistic Evidence from Hidden-Channel Distillation

arXiv:2607. 04432v1 Announce Type: cross Abstract: A student model trained on pure uniform noise can still inherit its teacher's digit-classification ability, provided the two share initialization.

arXiv Machine Learning
5d ago

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Direct feedback alignment (DFA) trains hidden layers via fixed random projections of output error, but with tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near a constant predictor of class frequencies. This stall is traced to the error’s common mode—a rank‑one component shared across inputs—that drives tanh units toward saturation. The study shows that calibration of the baseline readout to class priors suppresses collapse and speeds learning, while other interventions such as using Adam, adjusting feedback strength, or subtracting batch means affect the severity and recovery of collapse across MNIST, CIFAR‑10, and deeper networks.

By Varun Reddy, Bernardo L. Sabatini, Houman Safaai
arXiv Machine Learning
Sep 2

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

The paper investigates subliminal learning, where hidden traits from a teacher model are transferred to a student during distillation. It introduces trait‑direction drift as the underlying mechanism, showing that biased generation creates measurable preference gaps that accumulate into behavioral transfer during fine‑tuning. The authors propose probe‑space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction, significantly reducing hidden‑trait transfer while maintaining task performance.

By Zhixuan Liu, Zhichen Dong, Yuyu Fan, Xiangtian Li, Chao Yang
arXiv Machine Learning
1d ago

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

The paper investigates how teacher signals influence parameter updates in Multi‑Teacher On‑Policy Distillation (MOPD) by analyzing Qwen3‑1.7B and SmolLM3‑3B. It shows that loss averaging, Adam’s first‑moment bias, BF16 rounding, and the choice of averaging rule all shape the gradients and ultimately affect task performance. The study quantifies these effects, revealing, for example, that token‑averaging favors longer responses and that BF16 rounding masks most weight changes.

By Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang, Zhanyang Jin, Yihang Sun, Jiaxuan You
arXiv AI
3d ago

Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention

The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.

By Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
arXiv Computer Vision
Sep 3

Breaking the Geometric Bottleneck: Contrastive Expansion in Asymmetric Cross-Modal Distillation

The paper investigates how knowledge distillation from Vision Transformers to smaller CNNs can cause dimensional collapse in the student’s representation space. Using SVD and Shannon entropy, the authors show that cosine‑based distillation leads to a drastic reduction in effective rank, while adding an InfoNCE objective can double the rank but harms downstream accuracy due to signal dilution. They further demonstrate that a label‑aware contrastive objective (Supervised Contrastive distillation) can maintain or improve accuracy without unnecessary rank expansion, indicating that effective rank alone is not a reliable indicator of representation quality.

By Kabir Thayani
arXiv Machine Learning
Aug 28

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

The paper investigates how scaling the amount of model-generated, off‑task distillation data affects the recoverability of teacher‑induced traits in student models. In a controlled subliminal‑learning setup, teachers are prompted to express a target trait, producing restricted data such as number‑only completions. Students trained on larger independent datasets show a clearer manifestation of the teacher’s trait in a separate evaluation domain, with the effect being strongest when the trait is already favored or when alternative traits are present. The authors find this trend holds across model families, trait types, multi‑trait settings, and cross‑model transfer, and suggest that scaling should be coupled with trait‑aware curation and evaluation.

By Zhichen Dong, Zhixuan Liu, Yuyu Fan, Xiangtian Li, Shuyang Zhang, Chao Yang