arXiv:2501. 18530v3 Announce Type: replace-cross Abstract: We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional.
By Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
arXiv:2602. 10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets.
By Yuxin Ma, Nan Chen, Mateo D\'iaz, Soufiane Hayou, Dmitriy Kunisky, Soledad Villar
The paper introduces GFD-OPD, a method for on‑policy distillation of diffusion models that addresses challenges when compressing large teachers into smaller students. It identifies that standard distillation fails due to distribution gaps and classifier‑free guidance amplification, and proposes Fixed‑State KL to measure these gaps. GFD‑OPD reduces the student‑teacher discrepancy and achieves state‑of‑the‑art performance across multiple benchmarks.
By Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang, Jiazheng Xu, Wendi Zheng, Jie Tang, Dan Guo, Meng Wang
arXiv:2505. 24849v2 Announce Type: replace-cross Abstract: For three decades statistical mechanics has been providing a framework to analyse neural networks.
By Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
arXiv:2606. 29158v1 Announce Type: cross Abstract: Learning-rate transfer can reduce the cost of training large language models: instead of sweeping learning rates at target scale, practitioners extrapolate from smaller runs.
By Zaiwen Yang, Huaqing Zhang, Jing Xu, Jingzhao Zhang
arXiv:2607. 18026v1 Announce Type: new Abstract: Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment?
By Jiahe Fan, Yinghao Hou, Si Chen, Aiyuan Zhang, Hong Xie, Defu Lian