MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization
arXiv:2606. 17526v1 Announce Type: new Abstract: Efficient optimization is essential for training large language models.
arXiv:2606. 17526v1 Announce Type: new Abstract: Efficient optimization is essential for training large language models.
FedLore introduces a communication- and memory-efficient federated learning framework that shares a low-rank optimization basis across clients each round, mitigating subspace fragmentation and enabling exact low-rank aggregation. By refreshing this shared basis across rounds, FedLore allows model updates to exceed the per-round rank budget while maintaining a provable $O(T^{-1/2})$ stationarity bound under standard assumptions. Experiments on vision and language tasks, including federated pre‑training, demonstrate that FedLore outperforms low‑rank adapter baselines and matches or surpasses full‑parameter training while reducing communication and optimizer‑state memory.
arXiv:2511. 19716v3 Announce Type: replace-cross Abstract: Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise.
arXiv:2609.07666v1 Announce Type: new Abstract: Full-parameter fine-tuning of large language models has substantial memory costs because backpropagation stores activations and gradients. Zeroth-order...
arXiv:2607. 22906v1 Announce Type: new Abstract: We study adaptive gradient descent for continuously differentiable, possibly nonconvex objectives under one-sided H\"older regularity.
arXiv:2504. 12742v2 Announce Type: replace Abstract: Decentralized Federated Learning (DFL) enables collaborative model training without relying on a central server.
arXiv:2512. 23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollout policy $\pi_{\text{roll}}$.
arXiv:2602. 04396v2 Announce Type: replace-cross Abstract: Distributed training of foundation models via $\texttt{DDP}$ is limited by interconnect bandwidth.
StoSignSGD is a new sign‑based optimization algorithm that injects structural stochasticity into the sign operator, ensuring unbiased updates. It resolves the divergence issues of traditional SignSGD on non‑smooth objectives, achieving optimal convergence rates in convex settings and improved complexity bounds in non‑convex, non‑smooth problems. Empirical results show that StoSignSGD is stable and efficient across large language model training, outperforming AdamW and SignSGD in low‑precision regimes (FP8 and FP4) and delivering speedups and accuracy gains on models ranging from OLMo2‑370M to 7B LLMs.
arXiv:2606. 13657v2 Announce Type: replace Abstract: On-policy distillation (\textsc{OPD}) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student trajectories and dense teacher supervision.
The paper introduces Activation-Keyed Momentum (AK‑Momentum), a momentum update that uses the input activation of a linear layer as a key to apply a delta‑rule update, allowing each direction to decay at a rate proportional to its frequency of appearance. AK‑Momentum is proven to be a valid momentum, incorporates input‑side curvature correction without matrix inversion, and clears stale directions faster than traditional exponential moving average (EMA) under both fixed and drifting optima. It can replace the momentum buffer of any optimizer, scales with width under μP, adds only 22–25% extra compute, and demonstrates significant step‑count reductions in FineWeb‑Edu pretraining and other benchmarks. whyItMatters":"AK‑Momentum offers a principled, efficient way to adapt momentum decay to anisotropic training dynamics, improving convergence speed and stability across a range of models and datasets."
arXiv:2608. 06563v1 Announce Type: new Abstract: Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications.