arXiv Machine Learning By Dongmin Lee, William Lu, Anuran Makur

On the Oracle Complexity of Interpolation-Based Gradient Descent

Read the original on arXiv Machine Learning →

arXiv:2606. 19878v1 Announce Type: new Abstract: Recent work on first-order optimizers for empirical risk minimization (ERM) has suggested that smoothness of ERM loss functions in the training data, rather than in the optimization parameters, can be leveraged to improve the oracle complexity of gradient descent (GD) methods.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.