Towards Data Science

Why Gradient Descent Became Stochastic

A step-by-step journey from calculus-based optimization to Stochastic Gradient Descent The post Why Gradient Descent Became Stochastic appeared first on Towards Data Science .

arXiv Machine Learning
Jun 18

Stochastic Adaptive Gradient Descent Without Descent

arXiv:2509. 14969v2 Announce Type: replace Abstract: We introduce a new adaptive step-size strategy for convex optimization with stochastic gradient that exploits the local geometry of the objective function only by means of a first-order stochastic oracle and without any hyper-parameter tuning.

By Jean-Fran\c{c}ois Aujol, J\'er\'emie Bigot, Camille Castera
arXiv Machine Learning
Sep 15

Stochastic Gradient Descent over P2

The paper develops a diffusion approximation for stochastic gradient descent (SGD) when the optimization target is a functional on the Wasserstein space ℝ2. By lifting the problem to a Hilbert space via Lions differentiability, the authors construct a Gaussian random-field approximation whose velocity field matches the mean and covariance of the original stochastic gradient. They prove that this Gaussian approximation achieves second‑order weak accuracy, providing a rigorous basis for replacing sample‑driven randomness with analytically tractable Gaussian fluctuations in stochastic optimization over probability measures.

By Maria Oprea, Qin Li, Yunan Yang
Towards Data Science
Aug 27

The Sigmoid Function: From 'e' to Neural Networks

The article titled "The Sigmoid Function: From 'e' to Neural Networks" explores the origins and applications of the sigmoid function, a mathematical equation frequently used in data science and machine learning. It traces the function’s development from its foundational exponential form to its modern role in neural network architectures. The piece highlights how this simple yet powerful equation underpins many computational models in the field.

By Nikhil Dasari