arXiv:2609.38011v1 Announce Type: new
Abstract: Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downst...
By Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli
arXiv:2510. 12249v2 Announce Type: replace Abstract: In performative learning, the data distribution reacts to the deployed model - for example, because strategic users adapt their features to game it - which creates a more complex dynamic than in classical supervised learning.
By Edwige Cyffers, Alireza Mirrokni, Marco Mondelli
arXiv:2603.02069v2 Announce Type: replace
Abstract: We study scaling laws of signSGD under a power-law random features (PLRF) model that accounts for both feature and target decay. We analyze the pop...
By Jihwan Kim, Dogyoon Song, Chulhee Yun
arXiv:2607. 15450v1 Announce Type: cross Abstract: Self-distillation (SD) is typically studied when the student is retrained on the teacher's original training inputs.
By Hien Dang, Pratik Patil, Alessandro Rinaldo
arXiv:2601. 19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting.
By Mingyue Xu, Gal Vardi, Itay Safran
Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularizatio...