arXiv:2607. 22004v1 Announce Type: new Abstract: Energy natural gradient descent (ENGD) aligns parameter updates with the curvature of an underlying function-space energy, but existing formulations assume an unconstrained Euclidean parameter domain.
By Zhangyong Liang, Huanhuan Gao
arXiv:2608. 06218v1 Announce Type: cross Abstract: We study Muon, a recently proposed matrix-aware optimization method, in the context of the Stiefel manifold.
By Mikhail Solonko, Molozhavenko Alexander, Maxim Rakhuba
arXiv:2608. 12665v1 Announce Type: cross Abstract: For solving nonconvex equality-constrained optimization problems, a recent Gradient-Eigenstep Algorithm by Goyens et al.
By Frank E. Curtis, Lingjun Guo, Daniel P. Robinson
arXiv:2607. 07206v1 Announce Type: new Abstract: Adaptive optimizers mix several mechanisms: a metric or preconditioner maps gradients to descent directions, while estimation, memory, step-size control, constraints, stochasticity, target modification, and discretization determine which directions are available and how they are used.
By Zavier Li
arXiv:2606. 19411v1 Announce Type: new Abstract: Selecting a small, diverse, high-quality subset from a massive pool of candidates is a recurring primitive in modern machine learning -- data curation and coreset selection for training and fine-tuning large models, active-learning batch acquisition, prompt and exemplar selection for in-context learning, retrieval diversification, and experimental design.
By Richard Yi Da Xu
arXiv:2607. 07204v1 Announce Type: cross Abstract: Optimization geometrodynamics views optimizer state as evolving geometry.
By Zavier Li