Regularity-Aware Stochastic MGDA with Adaptive Conflict-Avoidant Update Direction Control
arXiv:2607. 15412v1 Announce Type: new Abstract: Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously.
arXiv:2608. 11749v1 Announce Type: cross Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation.
arXiv:2607. 15412v1 Announce Type: new Abstract: Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously.
arXiv:2608. 12704v1 Announce Type: cross Abstract: Multi-objective bilevel optimization has wide applications in the AI area such as automated learning and multi-task meta-learning.
arXiv:2609.07597v1 Announce Type: cross Abstract: Muon can be interpreted as optimizing a linear local objective over a spectral-norm ball. This gives a matrix-sign update that preserves the singular...
arXiv:2609. 30501v1 Announce Type: new Abstract: Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally convex (LLGC) settings (i.
SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning proposes a new scalarization method for multi-task learning that is invariant to the relative scales of task losses. By using a logarithmic transformation, SIMS converts the multi-objective problem into a single objective that preserves weak Pareto optimality and allows a smooth surrogate with controllable approximation error. Experiments on standard multi-task benchmarks show that SIMS consistently outperforms existing scalarization methods and achieves state‑of‑the‑art performance.
arXiv:2606. 03904v1 Announce Type: new Abstract: Multi-objective optimization (MOO) underlies many machine learning problems, yet MOO solvers across the loss-balancing, gradient-balancing, and Pareto-based families almost universally hand their reconciled directions to Adam~\cite{kingma2015adam}.
arXiv:2606. 15702v1 Announce Type: cross Abstract: Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm design and theoretical analysis.
arXiv:2405. 00914v4 Announce Type: replace-cross Abstract: We present in this paper novel accelerated fully first-order methods in \emph{Bilevel Optimization} (BLO).
arXiv:2606. 08783v1 Announce Type: cross Abstract: Orthogonalized momentum updates, as used in Muon-style optimizers, have recently shown strong empirical stability in large-scale deep learning.
arXiv:2605. 18106v3 Announce Type: replace-cross Abstract: A striking geometric disparity has long persisted in the practice of deep learning.
arXiv:2607. 22906v1 Announce Type: new Abstract: We study adaptive gradient descent for continuously differentiable, possibly nonconvex objectives under one-sided H\"older regularity.
The paper introduces a deterministic Simulated Oracle Direction (SOD) framework that enables escaping spurious local minima in non‑convex low‑rank matrix sensing without explicit tensor lifting. By projecting over‑parameterized escape directions back into the original parameter space, the method guarantees a strict decrease in objective value from existing local minima. Experiments show reliable escape from local minima and convergence to global optima with minimal computational overhead compared to explicit over‑parameterization.