Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models?
arXiv:2510. 24616v4 Announce Type: replace-cross Abstract: For four decades statistical physics has been providing a framework to analyse neural networks.
By Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
arXiv:2606. 05863v1 Announce Type: new Abstract: Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales.
By Hu Tan, Kuo Gai, Shihua Zhang
arXiv:2606. 21593v2 Announce Type: replace Abstract: Deep neural networks transform input data into latent representations that support a wide range of downstream tasks.
By Linara Adilova, Henning Petzka, Asja Fischer, Bernhard C. Geiger
arXiv:2512. 21315v2 Announce Type: replace Abstract: The data processing inequality is an information-theoretic principle stating that the information content of a signal cannot be increased by processing the observations.
By Roy Turgeman, Tom Tirer
arXiv:2606. 09658v1 Announce Type: cross Abstract: Muon has recently emerged as a state-of-the-art optimizer for pretraining Large Language Models (LLMs) and vision classifiers.
By Tianyu Ruan, Fengzhuo Zhang, Shuche Wang, Shihua Zhang