arXiv:2209. 03282v5 Announce Type: replace-cross Abstract: Accelerating the convergence of second-order optimization, particularly Newton-type methods, remains a pivotal challenge in algorithmic research.
By John Chiang
arXiv:2409. 03682v2 Announce Type: replace Abstract: Learning new tasks by leveraging prior experience is a fundamental trait of intelligent systems.
By El Mahdi Chayti, Martin Jaggi
arXiv:2607. 26562v1 Announce Type: cross Abstract: We study optimization under performative prediction, where deploying a model affects the future data distribution.
By Hiroki Hamaguchi, Yuya Hikima, Hiroshi Sawada, Akiko Takeda
arXiv:2608. 09523v1 Announce Type: new Abstract: Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success.
By Binchuan Qi
arXiv:2402.11215v4 Announce Type: replace
Abstract: The choice of batch size in minibatch stochastic gradient optimization is critical for both optimization and generalization performance in large-sc...
By Tim Tsz-Kit Lau, Han Liu, Mladen Kolar
The paper introduces ZFO, a lightweight framework that separates direction selection from step-size determination in large‑scale neural network optimization. ZFO uses a trusted first‑order optimizer to pick a search direction and then performs only two additional objective evaluations to build a local curvature‑aware model, selecting an adaptive step within a bounded interval. The authors provide theoretical guarantees for reliable curvature estimation, near‑optimal step selection, and convergence to a stationary point, and demonstrate that ZFO improves optimization and final performance over fixed‑step first‑order baselines on language‑model fine‑tuning tasks.
By Cristian McGee, El Houcine Bergou, Aritra Dutta