arXiv Machine Learning

Performance Evaluation of Ising and QUBO Variable Encodings in Boltzmann Machine Learning

arXiv Statistics ML
6d ago

Exact information accounting for SGD methods

The paper presents an exact, information‑theoretic analysis of stochastic gradient descent (SGD) and its variants, showing that a preconditioned SGD step corresponds to a posterior‑mean update in a Gaussian Bayes model. It decomposes one‑step regret into an intrinsic‑time cost and a change in comparator information, extending this split to an identity for the objective itself. The framework links convex convergence, saddle‑point escape, flatness‑generalization trade‑offs, learning‑rate schedules, adaptive optimizers, and various SGD variants, and it is validated on synthetic and real training runs, revealing how different optimizers achieve the same training loss through distinct step characteristics.

By Akshay Balsubramani
Hugging Face Trending Papers
Aug 2

Perspectives on Tsallis Statistics for Artificial Intelligence

Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter $q$ that controls the weight assigned to rare and frequent events. Originally proposed to describe physical systems with long-range correlations, multifractal geometry, and heavy-tailed fluctuations, the framework has become a recurring ingredient in modern artificial intelligence (AI): it underlies sparse attention mechanisms (\textsc{sparsemax} and $α$-\textsc{entmax}), maximum-entropy reinforcement learning with controllable exploration, robust and heavy-tailed probabilistic models, and a family of generalized loss functions and regularizers.

arXiv AI
6d ago

The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization

The paper investigates how the choice of reconstruction loss, specifically mean squared error (MSE), affects optimization in sequential post‑training quantization (PTQ) of large language models (LLMs). It identifies an Optimization Imbalance where MSE causes reconstruction loss magnitudes—and thus gradient magnitudes—to vary dramatically across quantization stages, leading to uneven parameter updates. The authors propose that root mean squared error (RMSE) variants, which implicitly normalize gradients, can decouple optimization strength from loss scale and outperform MSE as a drop‑in replacement.

By Chao Li, Shigeng Wang, Anbang Yao
arXiv Machine Learning
Jul 7

Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses

arXiv:2406. 14340v2 Announce Type: replace-cross Abstract: The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to zero (particularly, in the situation of constant learning rates).

By Steffen Dereich, Arnulf Jentzen, Adrian Riekert
arXiv AI
Jun 15

Gefen: Optimized Stochastic Optimizer

arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.

By Nadav Benedek, Tomer Koren, Ohad Fried