arXiv Machine Learning

Negative Stepsizes Make Gradient-Descent-Ascent Converge

arXiv:2505. 01423v2 Announce Type: replace-cross Abstract: Efficient computation of min-max problems is a central question in optimization, learning, games, and control.

arXiv Machine Learning
Aug 12

A lower bound for stepsize-based acceleration of gradient descent

arXiv:2608. 10418v1 Announce Type: cross Abstract: Recent work has shown that, for smooth convex optimization, plain gradient descent can be accelerated from its textbook convergence rate of $O(T^{-1})$ (where $T$ denotes the number of iterations) to $O\big(T^{-\log_2(1+\sqrt{2})}\big)$ using carefully designed stepsize schedules alone, without resorting to momentum or other algorithmic modifications.

By Jianhao Ma, Yuxin Chen
arXiv Machine Learning
Sep 10

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

The paper derives an exact discrete‑time law that captures how learning‑rate schedules and weight decay interact in scale‑invariant neural networks, showing that a single scalar quantity governs the effective step size. It demonstrates that the balance point between contraction and expansion is intrinsically unstable, leading to recurrent dynamics when using constant learning rates with weight decay. The authors extend this analysis to various optimizers and datasets, confirming the law’s precision and showing that performance peaks sharply at the predicted boundary.

By Hasan Amin, Wei-Kai Chang, Rajiv Khanna
arXiv AI
Sep 10

HyCO: A Hybrid Neural Solver for Combinatorial Optimization

HyCO is a hybrid neural solver that combines a sequential reinforcement learning (RL) solver with a global diffusion model (DM) to tackle combinatorial optimization problems. The RL component builds an initial solution prefix, after which HyCO switches to a conditional DM to finish the remaining decisions. The authors provide a theoretical framework showing that this hybrid approach yields lower expected regret than either method alone, identify an optimal trigger step for the switch, and implement a lightweight adaptive trigger based on policy entropy and RL‑DM disagreement, achieving consistent performance gains across benchmarks.

By Yuheng Li, Di Yang, Haipeng Chen, Yanhai Xiong
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas
Hugging Face Trending Papers
Sep 8

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

The paper investigates how normalization makes neural networks scale‑invariant, creating a feedback loop between learning‑rate schedules and weight decay that controls the effective step size of the optimizer. It derives an exact discrete‑time law showing that a single scalar quantity captures all schedule and decay effects, with norm growth providing a self‑quenching counter‑force that defines a sharp boundary between contraction‑ and expansion‑dominated regimes. Through exact analysis of a normalized regression model and experiments on MLPs, CNNs, GPT‑2, and various datasets, the authors demonstrate that constant learning rates with weight decay are intrinsically unstable, leading to recurrent dynamics, and that adaptive optimizers exhibit weaker stabilization under normalization. "whyItMatters":"The study provides a precise, actionable rule for controlling training dynamics and schedule design in modern deep learning by isolating a single governing quantity for scale‑invariant optimization."

arXiv Machine Learning
1d ago

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

The paper introduces ZFO, a lightweight framework that separates direction selection from step-size determination in large‑scale neural network optimization. ZFO uses a trusted first‑order optimizer to pick a search direction and then performs only two additional objective evaluations to build a local curvature‑aware model, selecting an adaptive step within a bounded interval. The authors provide theoretical guarantees for reliable curvature estimation, near‑optimal step selection, and convergence to a stationary point, and demonstrate that ZFO improves optimization and final performance over fixed‑step first‑order baselines on language‑model fine‑tuning tasks.

By Cristian McGee, El Houcine Bergou, Aritra Dutta