arXiv Machine Learning By El Mahdi Chayti, Martin Jaggi

A New First-Order Meta-Learning Algorithm with Convergence Guarantees

Read the original on arXiv Machine Learning →

arXiv:2409. 03682v2 Announce Type: replace Abstract: Learning new tasks by leveraging prior experience is a fundamental trait of intelligent systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 27

Greedy dynamical meta-learning

Gradient descent scales well to large models, but becomes unstable over long time horizons. Gradient-free optimizers can scale to arbitrary timespans, but are hobbled by high dimensions.

arXiv Machine Learning
Sep 22

Personalized Federated Reinforcement Learning via Model-Agnostic Meta-Learning: Convergence of Exact and Hessian-Free Meta-Policy Gradients

arXiv:2609. 22833v1 Announce Type: new Abstract: We study personalized federated reinforcement learning, in which $n$ agents, each acting in its own Markov decision process, collaborate through a server to learn a shared MAML-style policy initialization that becomes effective for an individual agent once that agent adapts it with a single local policy-gradient step.

By Ali Beikmohammadi, Sarit Khirirat, Sindri Magn\'usson
arXiv Machine Learning
1d ago

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

The paper introduces ZFO, a lightweight framework that separates direction selection from step-size determination in large‑scale neural network optimization. ZFO uses a trusted first‑order optimizer to pick a search direction and then performs only two additional objective evaluations to build a local curvature‑aware model, selecting an adaptive step within a bounded interval. The authors provide theoretical guarantees for reliable curvature estimation, near‑optimal step selection, and convergence to a stationary point, and demonstrate that ZFO improves optimization and final performance over fixed‑step first‑order baselines on language‑model fine‑tuning tasks.

By Cristian McGee, El Houcine Bergou, Aritra Dutta