arXiv:2606. 10913v1 Announce Type: new Abstract: We explore whether intrinsic symmetries of the training data lead to conserved quantities during gradient-flow training of neural networks.
By Jakob Galley, Vahid Shahverdi, Axel Flinth
arXiv:2606. 11341v1 Announce Type: new Abstract: Modular neural network pipelines suffer from error compounding: noise at any module boundary propagates and potentially amplifies through subsequent modules.
By David Young, Swan Yi Htet
arXiv:2606. 09744v1 Announce Type: new Abstract: We study feed-forward ReLU networks with fixed readout and quadratic loss.
By Claudio Nordio
arXiv:2501. 02436v5 Announce Type: replace Abstract: Advancements in artificial intelligence call for a deeper understanding of the fundamental mechanisms underlying deep learning.
By Yuchen Lin, Yong Zhang, Sihan Feng, Hong Zhao
arXiv:2501. 07400v2 Announce Type: replace-cross Abstract: We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the input layer, and under the assumption that the weights are, in a precise sense, adapted to the coordinate system distinguished by the activations.
By Thomas Chen
arXiv:2608. 04442v1 Announce Type: new Abstract: Robustness to natural corruptions remains a fundamental challenge for deep neural networks.
By Jiangang Yang, Wenhui Shi, Lu Hu, Jing Xing, Jian Liu