Lightweight and Versatile Learned Optimization by Recombination of Gradient History
Read the original on arXiv Machine Learning →The paper introduces a lightweight learned optimizer that recombines gradient history by averaging over disjoint time spans, reducing the prediction space to a single scalar coefficient per average shared across parameters. By progressively averaging older gradients, the method keeps memory usage low while maintaining independent contributions from long‑term history. A small 37k‑parameter network trained in under a GPU‑hour generalizes zero‑shot to unseen tasks, improving validation loss on BERT‑Tiny and GPT‑Tiny and boosting test accuracy over Adam on Vision Transformers and graph models with minimal FLOPs overhead.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.