arXiv:2609.15679v1 Announce Type: cross
Abstract: This paper studies projection-free algorithms for stochastic constrained multi-level compositional optimization. In this context, the objective funct...
By Wei Jiang, Sifan Yang, Wenhao Yang, Yibo Wang, Yuanyu Wan, Zechao Li, Lijun Zhang
arXiv:2406. 13041v3 Announce Type: replace Abstract: Lower-bound analyses for nonconvex strongly-concave minimax optimization problems have shown that stochastic first-order algorithms require at least $\mathcal{O}(\varepsilon^{-4})$ sample complexity to find an $\varepsilon$-stationary point.
By Haoyuan Cai, Sulaiman A. Alghunaim, Ali H. Sayed
arXiv:2607. 08104v1 Announce Type: new Abstract: Stochastic gradient descent (SGD) is a cornerstone of modern optimization.
By Ryusei Yamada, Naoki Sato, Hideaki Iiduka
arXiv:2510.11676v2 Announce Type: replace-cross
Abstract: We study convex composite optimization problems, where the objective function is given by the sum of a prox-friendly function and a convex fu...
By Chuan He, Bowen Li, Zhaosong Lu
arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.
By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
arXiv:2609.15723v1 Announce Type: new
Abstract: Traditional variance reduction methods (e.g., SPIDER, SARAH, STORM) have been extensively investigated for improving the convergence rates of stochasti...
By Wei Jiang, Sifan Yang, Yibo Wang, Lijun Zhang, Zechao Li
arXiv:2608. 12009v1 Announce Type: cross Abstract: Bregman proximal stochastic gradient (BPSG) methods bring variance-reduced composite optimization to objectives whose geometry is poorly captured by Euclidean smoothness.
By Chenhan Jin, Shengze Xu, Binghui Xie, Kaiwen Zhou, Fan Jia, James Cheng, Tieyong Zeng
arXiv:2505. 01258v2 Announce Type: replace-cross Abstract: Bilevel optimization has recently attracted significant attention in machine learning due to its wide range of applications and advanced hierarchical optimization capabilities.
By Tianshu Chu, Dachuan Xu, Wei Yao, Chengming Yu, Jin Zhang
arXiv:2605. 26000v2 Announce Type: replace-cross Abstract: Stochastic gradient descent (SGD) is foundational to large-scale statistical learning and stochastic optimization.
By Jose Blanchet, Peter Glynn, Wenhao Yang
arXiv:2605.28517v2 Announce Type: replace-cross
Abstract: Stochastic gradient descent with momentum (SGDM) is one of the most widely used optimization algorithms in machine learning. While optimizati...
By Yunwen Lei, Zimeng Wang, Xiaoming Yuan
The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.
By Feng Zhu, Robert W. Heath Jr., Aritra Mitra
arXiv:2607. 14731v1 Announce Type: new Abstract: Local SGD, also known as Federated Averaging, is a widely used distributed optimization algorithm.
By Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich, Aurelien Lucchi, Eduard Gorbunov, Lingxiao Wang