arXiv:2607. 06151v1 Announce Type: new Abstract: Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitting and reduced performance on unseen data.
By Yao Fu, Chunxia Zhang, Junmin Liu, Yihang Jin, Haishan Ye, Yuanao Yang
arXiv:2607. 16261v1 Announce Type: cross Abstract: Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments.
By Apostolos Avranas
arXiv:2501. 10870v2 Announce Type: replace-cross Abstract: The principal objective of this work is twofold within nonparametric regression settings: (1) to establish the minimax optimal convergence rates for fixed-bandwidth Gaussian kernel spectral algorithms when the true regression function resides in a Sobolev space, and (2) to apply Gaussian spectral algorithms for achieving robust and adaptive transfer learning under concept shift.
By Haotian Lin, Matthew Reimherr
arXiv:2606. 16050v1 Announce Type: cross Abstract: Robust deep learning under heavy-tailed and impulsive noise remains challenging because conventional losses such as mean squared error (MSE) exhibit unbounded sensitivity to outliers.
By Mainak Kundu, Ria Kanjilal, Ismail Uysal
arXiv:2606. 17567v1 Announce Type: new Abstract: While sequential residual fitting is the bedrock of standard boosting frameworks, it inherently breeds learner redundancy by repeatedly revisiting correlated error components.
By Ye Su, Jipeng Guo, Yong Liu, Xin Xu, Gangchun Zhang, Jinxin Chen, Di Wu, Longlong Zhao
arXiv:2605. 27756v2 Announce Type: replace-cross Abstract: Linear dimensionality reduction methods such as proper orthogonal decomposition (POD) make high-dimensional data amenable to analysis by identifying the principal components, or modes, that capture the most variance, or energy, in the data and constructing a low-dimensional representation in the subspace they span.
By Tomoki Koike, Prakash Mohan, Marc T. Henry de Frahan, Elizabeth Qian, Julie Bessac
arXiv:2605. 29547v2 Announce Type: replace-cross Abstract: Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators.
By Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang
The paper introduces Wave-BLS, a robust Broad Learning System that replaces the traditional squared error loss with an asymmetric, bounded, and smooth wave loss function. This change allows controlled penalization of large errors and eliminates the need for matrix inversion by using a Nesterov accelerated gradient scheme. Experiments on 30 UCI datasets show that Wave-BLS consistently outperforms classical BLS and other robust variants, with statistical tests confirming the significance of the improvements and demonstrating greater resilience to noise and outliers.
By Mushir Akhtar, A. Varshney, A. Quadir, A. Rahaman, M. Tanveer, Mohd. Arshad
arXiv:2606. 23942v1 Announce Type: new Abstract: We present a large-scale empirical study isolating the contributions of the Derivative Regularization penalty (DREG).
By Rowan Martnishn
arXiv:2609.21017v1 Announce Type: cross
Abstract: We study data-driven early stopping for spectral regularisation methods in the classical non-parametric regression setting. Building on the discrepan...
By Mike Nguyen, Nicole M\"ucke
arXiv:2607. 16261v2 Announce Type: replace-cross Abstract: Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments.
By Apostolos Avranas
arXiv:2607. 05201v1 Announce Type: new Abstract: In non-stationary streaming environments, simultaneously adapting to complex, non-linear domain shifts via continual learning while mitigating the catastrophic effects of severe, uncalibrated label noise poses a fundamental mathematical challenge.
By Rai Hisada, Kanji Tanaka