arXiv AI By Srijan Tiwari, Aditya Chauhan, Manjot Singh

Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization

Read the original on arXiv AI →

arXiv:2606. 32000v1 Announce Type: cross Abstract: Why do neural networks memorize algorithmic training data long before they generalize?

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 27

Hyperball May Not Be a Free Lunch

arXiv:2607. 22444v1 Announce Type: new Abstract: For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates.

By Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai