arXiv AI By Srijan Tiwari, Aditya Chauhan, Manjot Singh

Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization

Read the original on arXiv AI →

arXiv:2606. 32000v1 Announce Type: cross Abstract: Why do neural networks memorize algorithmic training data long before they generalize?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 27

Hyperball May Not Be a Free Lunch

arXiv:2607. 22444v1 Announce Type: new Abstract: For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates.

By Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai