Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks
arXiv:2502. 11152v4 Announce Type: replace-cross Abstract: The optimization foundations of deep linear networks have recently received significant attention.
arXiv:2403. 10232v2 Announce Type: replace-cross Abstract: Conventional matrix completion methods approximate the missing values by assuming the matrix to be low-rank, which leads to a linear approximation of missing values.
arXiv:2502. 11152v4 Announce Type: replace-cross Abstract: The optimization foundations of deep linear networks have recently received significant attention.
The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.
arXiv:2506. 13139v3 Announce Type: replace-cross Abstract: Modern Machine Learning (ML) and Deep Neural Networks (DNNs) often operate on high-dimensional data and rely on overparameterized models, where classical low-dimensional intuitions break down.
arXiv:2606. 29951v1 Announce Type: new Abstract: Interpretable Mesomorphic Neural Networks (IMNs) offer a promising framework that combines the predictive power of deep neural networks with the interpretability of linear models.
arXiv:2607. 07735v1 Announce Type: cross Abstract: Sparse precision matrix estimation provides an interpretable and computationally efficient framework for modeling conditional dependencies in high-dimensional, low-sample-size data.
The paper introduces a least‑squares method for training quadratic neural networks with regularization, providing a lower bound on the training optimization problem when the regularization coefficient is positive. It delivers closed‑form expressions for both the approximate solution and its sensitivity to data errors, and shows that the solution is optimal when the regularization coefficient is zero. The approach offers computational advantages over iterative methods like backpropagation and is validated on a nonlinear system identification example.
arXiv:2608. 09523v1 Announce Type: new Abstract: Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success.
arXiv:2602.01437v2 Announce Type: replace-cross Abstract: The problem of corrupted data, missing features, or missing modalities continues to plague the modern machine learning landscape. To address...
arXiv:2504.16450v4 Announce Type: replace Abstract: We derive a differential equation that governs the evolution of the generalization gap when a model is trained by gradient descent-based methods. T...
arXiv:2510. 01175v2 Announce Type: replace Abstract: While normalization techniques are widely used in deep learning, their theoretical understanding remains relatively limited.
arXiv:2606. 28662v1 Announce Type: cross Abstract: The flatness hypothesis suggests that flatness of the loss landscape, as measured by the eigenvalues of the loss Hessian, correlates with better neural network generalization.
arXiv:2601. 16884v3 Announce Type: replace Abstract: We study multigrade deep learning (MGDL) as a principled framework for structured error refinement in deep neural networks.