arXiv:2604. 00316v2 Announce Type: replace-cross Abstract: Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that.
By Marcel Tom\`as Bernal, Neil Rohit Mallinar, Mikhail Belkin
arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.
By Ege Erdogan, Ana Lucic
arXiv:2606. 09049v1 Announce Type: cross Abstract: We propose the data augmented bootstrap (DAB), a framework for constructing confidence intervals from approximately invariant transformations of the data.
By Kevin Han Huang
We propose the data augmented bootstrap (DAB), a framework for constructing confidence intervals from approximately invariant transformations of the data. As special cases, DAB recovers popular methods that rely on exact group symmetries, such as conformal prediction, wild bootstrap for Maximum Mean Discrepancy U-statistics and the recently proposed SymmPI.
arXiv:2505. 19809v3 Announce Type: replace-cross Abstract: In many real-world applications of regression, conditional probability estimation, and uncertainty quantification, exploiting symmetries rooted in physics or geometry can dramatically improve generalization and sample efficiency.
By Daniel Ordo\~nez-Apraez, Vladimir Kosti\'c, Alek Fr\"ohlich, Vivien Brandt, Karim Lounici, Massimiliano Pontil
arXiv:2406. 13944v2 Announce Type: replace-cross Abstract: This paper establishes the generalization error of pooled min-$\ell_2$-norm interpolation in transfer learning, where data from diverse distributions are available.
By Yanke Song, Kenneth Gu, Sohom Bhattacharya, Pragya Sur
arXiv:2510. 17303v2 Announce Type: replace Abstract: Symmetries are known to improve the empirical performance of machine learning models, yet theoretical guarantees explaining these gains remain limited.
By Armin Beck, Peter Ochs
arXiv:2603. 27631v2 Announce Type: replace Abstract: Self-supervised pre-training, where large corpora of unlabeled data are used to learn representations for downstream fine-tuning, has become a cornerstone of modern machine learning.
By Mohammad Tinati, Stephen Tu
arXiv:2606. 15832v1 Announce Type: new Abstract: Empirical risk minimization on massive datasets naturally exhibits a nested double finite-sum structure, where $N=nm$ total samples are logically or physically partitioned into $n$ blocks of size $m$ (e.
By Igor Sokolov, Laurent Condat, Peter Richt\'arik
arXiv:2510. 24616v4 Announce Type: replace-cross Abstract: For four decades statistical physics has been providing a framework to analyse neural networks.
By Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
arXiv:2207. 03116v4 Announce Type: replace Abstract: We introduce a general method for learning representations that are equivariant to symmetries of data.
By Giovanni Luca Marchetti, Gustaf Tegn\'er, Anastasiia Varava, Danica Kragic
arXiv:2512. 20043v3 Announce Type: replace Abstract: Symmetry is fundamental to understanding physical systems and can improve performance and sample efficiency in machine learning.
By Yuxuan Chen, Jung Yeon Park, Floor Eijkelboom, Jianke Yang, Jan-Willem van de Meent, Lawson L. S. Wong, Robin Walters