arXiv Machine Learning By Zijian Liu

Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad

Read the original on arXiv Machine Learning →

arXiv:2605. 18694v2 Announce Type: replace-cross Abstract: Many tasks in modern machine learning are observed to involve heavy-tailed gradient noise during the optimization process.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.