arXiv Machine Learning By Arzu Ahmadova, Ismail Huseynov

Learning from the Descent Direction: Adaptive Gradient Descent under One-Sided H\"older Regularity

Read the original on arXiv Machine Learning →

arXiv:2607. 22906v1 Announce Type: new Abstract: We study adaptive gradient descent for continuously differentiable, possibly nonconvex objectives under one-sided H\"older regularity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

The paper introduces ZFO, a lightweight framework that separates direction selection from step-size determination in large‑scale neural network optimization. ZFO uses a trusted first‑order optimizer to pick a search direction and then performs only two additional objective evaluations to build a local curvature‑aware model, selecting an adaptive step within a bounded interval. The authors provide theoretical guarantees for reliable curvature estimation, near‑optimal step selection, and convergence to a stationary point, and demonstrate that ZFO improves optimization and final performance over fixed‑step first‑order baselines on language‑model fine‑tuning tasks.

By Cristian McGee, El Houcine Bergou, Aritra Dutta