The paper presents an exact, information‑theoretic analysis of stochastic gradient descent (SGD) and its variants, showing that a preconditioned SGD step corresponds to a posterior‑mean update in a Gaussian Bayes model. It decomposes one‑step regret into an intrinsic‑time cost and a change in comparator information, extending this split to an identity for the objective itself. The framework links convex convergence, saddle‑point escape, flatness‑generalization trade‑offs, learning‑rate schedules, adaptive optimizers, and various SGD variants, and it is validated on synthetic and real training runs, revealing how different optimizers achieve the same training loss through distinct step characteristics.
By Akshay Balsubramani
Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter $q$ that controls the weight assigned to rare and frequent events. Originally proposed to describe physical systems with long-range correlations, multifractal geometry, and heavy-tailed fluctuations, the framework has become a recurring ingredient in modern artificial intelligence (AI): it underlies sparse attention mechanisms (\textsc{sparsemax} and $α$-\textsc{entmax}), maximum-entropy reinforcement learning with controllable exploration, robust and heavy-tailed probabilistic models, and a family of generalized loss functions and regularizers.
The paper investigates how the choice of reconstruction loss, specifically mean squared error (MSE), affects optimization in sequential post‑training quantization (PTQ) of large language models (LLMs). It identifies an Optimization Imbalance where MSE causes reconstruction loss magnitudes—and thus gradient magnitudes—to vary dramatically across quantization stages, leading to uneven parameter updates. The authors propose that root mean squared error (RMSE) variants, which implicitly normalize gradients, can decouple optimization strength from loss scale and outperform MSE as a drop‑in replacement.
By Chao Li, Shigeng Wang, Anbang Yao
arXiv:2406. 14340v2 Announce Type: replace-cross Abstract: The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to zero (particularly, in the situation of constant learning rates).
By Steffen Dereich, Arnulf Jentzen, Adrian Riekert
arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.
By Nadav Benedek, Tomer Koren, Ohad Fried
arXiv:2605. 27991v2 Announce Type: replace-cross Abstract: Gradient-flow optimization is usually viewed as an algorithmic procedure for minimizing empirical loss, with training duration selected by validation or heuristic early-stopping rules.
By Minhao Yao, Ruoyu Wang, Xihong Lin, Lin Liu, Zhonghua Liu
arXiv:2505. 11635v2 Announce Type: cross Abstract: Many real-world tasks, from associative memory to symbolic reasoning, benefit from discrete, structured representations that standard continuous latent models can struggle to express.
By Nikhil Kapasi, Mohamed Elfouly, William Whitehead, Luke Theogarajan
arXiv:2610.08612v1 Announce Type: cross
Abstract: We consider a modular associative neural network made of $L$ Hopfield models (HMs), coupled so that intra-module interactions are Hebbian and inter-m...
By Elena Agliari, Andrea Lepre, Edoardo Roscani
arXiv:2609.08219v1 Announce Type: cross
Abstract: Neural networks acquire internal representations through learning. In this work, we formulate stochastic gradient descent (SGD) as a Markovian stocha...
By Shuta Kobayashi, Andreas Dechant
arXiv:2511. 02496v2 Announce Type: replace Abstract: We study latent geometry as an explicit component of representation quality in data-scarce learning.
By Ronald Katende
arXiv:2607. 04993v1 Announce Type: cross Abstract: Many phenomena of deep learning are dynamical: they concern not only which minima exist, but how gradient descent reaches, avoids, or selects among them.
By Thomas Hofmann
arXiv:2609.15643v1 Announce Type: new
Abstract: Flow-matching diffusion models have recently emerged as a strong paradigm for high-fidelity visual generation. However, their prohibitively high fine-t...
By Jiayang Gu, Zheng Fang, Lichaun Xiang, Fanghui Liu, Xu Cai, Hongkai Wen