Convergence, design and training of continuous-time dropout as a random batch method
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper develops a diffusion approximation for stochastic gradient descent (SGD) when the optimization target is a functional on the Wasserstein space ℝ2. By lifting the problem to a Hilbert space via Lions differentiability, the authors construct a Gaussian random-field approximation whose velocity field matches the mean and covariance of the original stochastic gradient. They prove that this Gaussian approximation achieves second‑order weak accuracy, providing a rigorous basis for replacing sample‑driven randomness with analytically tractable Gaussian fluctuations in stochastic optimization over probability measures.
arXiv:2411. 01982v2 Announce Type: replace-cross Abstract: We study the problem of learning controlled stochastic differential equations (SDEs) \[ dX_t = b(t,X_t,u_t)\,dt + \sigma(t,X_t,u_t)\,dW_t, \] whose drift and diffusion depend nonlinearly on time, state, and control values.
arXiv:2607. 08757v1 Announce Type: cross Abstract: Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory.
arXiv:2603. 20467v2 Announce Type: replace-cross Abstract: Stochastic differential equations (SDEs), which serve as the governing equations for dynamical systems in a broad range of applications, can become cost-prohibitive for numerical simulation at scales necessary for quantifying key properties.
arXiv:2606. 13796v1 Announce Type: cross Abstract: Recursive training of generative models on their own outputs can lead to model collapse, a compounding drift away from the true data distribution.
arXiv:2602.13960v2 Announce Type: replace Abstract: Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is ty...