arXiv Machine Learning

Momentum LMS Theory beyond Stationarity: Stability, Tracking, and Regret

arXiv:2602. 11995v2 Announce Type: replace Abstract: In large-scale data processing scenarios, data often arrive in sequential streams generated by complex systems that exhibit drifting distributions and time-varying system parameters.

arXiv Machine Learning
Aug 26

Adaptive prediction theory combining offline and online learning

The paper studies a two‑stage learning framework that first trains an offline model using approximate nonlinear‑least‑squares estimation and then adapts it online with a meta‑LMS algorithm to handle parameter drift in nonlinear stochastic dynamical systems. It provides an upper bound on the offline generalization error that accounts for strong data correlation and distribution shift via Kullback‑Leibler divergence, and it demonstrates that the combined offline‑online approach outperforms methods that rely solely on offline or online learning. Both theoretical analysis and empirical experiments support the claimed performance gains.

By Haizheng Li, Lei Guo
arXiv Machine Learning
Sep 14

Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization

The paper presents a theoretical study of Adam in non‑stationary stochastic optimization, distinguishing two regimes: Euclidean tracking under adaptive strong monotonicity and high‑probability projected stationarity for general smooth objectives. It derives finite‑time bounds that decompose into initialization, objective drift, first‑moment tracking error (β₁), and preconditioner perturbation (β₂), and characterizes burn‑in times for constant and step‑decay schedules. The analysis reveals a noise–drift tradeoff, showing that in noise‑dominated settings Adam’s adaptive mechanisms can improve guarantees, while in drift‑dominated settings they may worsen tracking, potentially making vanilla SGD preferable.

By Sharan Sahu, Abir Sarkar, Cameron J. Hogan, Martin T. Wells
arXiv Machine Learning
Aug 18

Adaptive Optimization via Momentum on Variance-Normalized Gradients

arXiv:2602. 10204v2 Announce Type: replace Abstract: We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary ideas: variance-based normalization and momentum applied after normalization.

By Francisco Patitucci, Aryan Mokhtari
arXiv Machine Learning
5d ago

Online Learning via Learned Latent Bayesian Tracking

The paper introduces AURA, a meta‑learning framework that learns a low‑dimensional latent state‑space model for the evolution of optimal model parameters under distribution shift. Online adaptation is performed via extended Kalman filtering in this latent space, followed by reconstruction of full model parameters through a learned lifting map, enabling efficient single‑step updates. Experiments on neural wireless receivers and non‑stationary image classification show that AURA improves adaptation speed, accuracy, and computational efficiency compared to existing online learning and Bayesian filtering baselines.

By Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone