Performance Evaluation of Ising and QUBO Variable Encodings in Boltzmann Machine Learning
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper presents an exact, information‑theoretic analysis of stochastic gradient descent (SGD) and its variants, showing that a preconditioned SGD step corresponds to a posterior‑mean update in a Gaussian Bayes model. It decomposes one‑step regret into an intrinsic‑time cost and a change in comparator information, extending this split to an identity for the objective itself. The framework links convex convergence, saddle‑point escape, flatness‑generalization trade‑offs, learning‑rate schedules, adaptive optimizers, and various SGD variants, and it is validated on synthetic and real training runs, revealing how different optimizers achieve the same training loss through distinct step characteristics.
Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter $q$ that controls the weight assigned to rare and frequent events. Originally proposed to describe physical systems with long-range correlations, multifractal geometry, and heavy-tailed fluctuations, the framework has become a recurring ingredient in modern artificial intelligence (AI): it underlies sparse attention mechanisms (\textsc{sparsemax} and $α$-\textsc{entmax}), maximum-entropy reinforcement learning with controllable exploration, robust and heavy-tailed probabilistic models, and a family of generalized loss functions and regularizers.
The paper investigates how the choice of reconstruction loss, specifically mean squared error (MSE), affects optimization in sequential post‑training quantization (PTQ) of large language models (LLMs). It identifies an Optimization Imbalance where MSE causes reconstruction loss magnitudes—and thus gradient magnitudes—to vary dramatically across quantization stages, leading to uneven parameter updates. The authors propose that root mean squared error (RMSE) variants, which implicitly normalize gradients, can decouple optimization strength from loss scale and outperform MSE as a drop‑in replacement.
arXiv:2406. 14340v2 Announce Type: replace-cross Abstract: The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to zero (particularly, in the situation of constant learning rates).
arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.
arXiv:2605. 27991v2 Announce Type: replace-cross Abstract: Gradient-flow optimization is usually viewed as an algorithmic procedure for minimizing empirical loss, with training duration selected by validation or heuristic early-stopping rules.