arXiv:2606. 05219v1 Announce Type: new Abstract: Recent analyses of multi-pathway Deep Linear Networks use Gradient Flow to predict a "winner-takes-all" specialization in which path symmetry breaks and each feature concentrates in a single pathway.
By Hee-Sung Kim, Sungyoon Lee
arXiv:2604. 08742v2 Announce Type: replace-cross Abstract: Adam is widely used, but its convergence theory remains incomplete even in the deterministic full-batch setting because momentum and adaptive preconditioning are tightly coupled.
By Yaxin Yu, Long Chen, Zeyi Xu
arXiv:2602. 03024v2 Announce Type: replace-cross Abstract: Deep Equilibrium Models (DEQs) have emerged as a powerful paradigm in deep learning, offering the ability to model infinite-depth networks with constant memory usage.
By Junchao Lin, Zenan Ling, Jingwen Xu, Robert C. Qiu
arXiv:2606. 09112v1 Announce Type: cross Abstract: The rapid evolution of artificial intelligence has led to substantial advances in deep neural networks.
By Chen-Rui Fan, Bo Lu, Xing-Yu Wu, Tie-Jun Wang, Chuan Wang
arXiv:2607. 07425v1 Announce Type: cross Abstract: Many biological processes are governed by complex dynamical mechanisms that remain incompletely understood despite increasing volumes of experimental data.
By Rebecca M. Crossley, Yuan Yin, Sarah L. Waters, Ruth E. Baker
arXiv:2607. 24996v1 Announce Type: cross Abstract: Neural networks are hindered by accumulating dormant neurons and loss of expressivity throughout training, particularly in non-stationary data settings, such as continual supervised and reinforcement learning.
By Luc McCutcheon, Evangelos Chatzaroulas, Saber Fallah