Flatland: The Adventures of Gradient Descent with Large Step Sizes
arXiv:2606. 06722v1 Announce Type: new Abstract: The training of neural networks often entails objective functions that are not globally $L$-smooth.
arXiv:2507. 20424v3 Announce Type: replace Abstract: We study centralized distributed data parallel training of deep neural networks (DNNs), aiming to improve the trade-off between communication efficiency and model performance of the local gradient methods.
arXiv:2606. 06722v1 Announce Type: new Abstract: The training of neural networks often entails objective functions that are not globally $L$-smooth.
arXiv:2512. 12737v2 Announce Type: replace Abstract: Decentralized federated learning (DFL) enables collaborative model training without a central server, but converges slowly under statistical heterogeneity.
arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
arXiv:2608. 09523v1 Announce Type: new Abstract: Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success.
The paper compares distributed adversarial training algorithms—both centralized and decentralized—within multi‑agent learning environments. It introduces a theoretical framework to analyze how efficiently these algorithms escape local minima, a property linked to model flatness and robustness. The study finds that with small perturbation bounds and large batch sizes, decentralized methods (consensus and diffusion) escape local minima faster than centralized ones, but this advantage may diminish as attack strength increases.
arXiv:2502. 11152v4 Announce Type: replace-cross Abstract: The optimization foundations of deep linear networks have recently received significant attention.
arXiv:2608.24568v1 Announce Type: cross Abstract: Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phenomenon often associated with the geomet...
arXiv:2301. 06308v2 Announce Type: replace-cross Abstract: Sharpness-aware minimization (SAM) is a training method that seeks to find flat minima in deep learning, resulting in state-of-the-art performance across various domains.
arXiv:2502. 17055v5 Announce Type: replace Abstract: Training instability in modern deep learning systems is frequently triggered by rare but extreme gradient-norm spikes, which can induce oversized parameter updates, corrupt optimizer state, and lead to slow recovery or divergence.
The paper investigates why the SCAFFOLD algorithm, designed to be robust to data heterogeneity in federated learning, often underperforms compared to the simpler FedAvg. It identifies the presence of Edge of Stability (EoS) dynamics and progressive sharpening as key factors, showing that both algorithms exhibit EoS behavior across various architectures and hyperparameters. Crucially, the study finds that at the EoS, SCAFFOLD’s ability to estimate the global gradient deteriorates, as indicated by a weakened correlation between sharpness and gradient estimation error, explaining its limited practical advantage.
arXiv:2607. 06151v1 Announce Type: new Abstract: Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitting and reduced performance on unseen data.
AYLA is a loss reparameterization framework that applies a sigmoid‑controlled power‑law transformation to the empirical loss, dynamically adjusting gradient magnitudes without changing stationary points or optimal solutions. By reshaping optimization trajectories, AYLA accelerates descent in flat or saddle‑dominated regions and stabilizes late‑stage training, leading to improved feature recovery in two‑layer tanh networks on synthetic Gaussian data. Experiments show enhanced weight alignment, neuron similarity, activation correlation, and richer internal representations, while mitigating rank collapse and promoting a transition from lazy to active feature‑learning regimes.