arXiv:2606. 02857v1 Announce Type: cross Abstract: Zeroth-order (ZO) optimization is a memory-efficient alternative to backpropagation for fine-tuning large language models, but its deployment is limited by the high variance of gradient estimation.
By Liyan Tan, Yequan Zhao, Yifan Yang, Ruijie Zhang, Xinling Yu, Zheng Zhang
arXiv:2606. 14970v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) has become a central application of modern optimization, enabling pretrained models to adapt to diverse downstream tasks and domain-specific data.
By Dmitriy Bystrov, Daniil Medyakov, Dmitry Bylinkin, Aleksandr Beznosikov
arXiv:2606. 05516v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables memory-efficient fine-tuning of large language models (LLMs) using only forward passes, but it remains unclear how useful adaptation is distributed across layers.
By Wanhao Yu, Ziyan Wang, Zheng Wang, Abeer Matar Almalky, Yihang Zuo, Shuteng Niu, Sen Lin, Adnan Siraj Rakin, Deliang Fan, Li Yang
arXiv:2506. 21833v2 Announce Type: replace Abstract: Forward-mode automatic differentiation (FmAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagation-free alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting memory-efficient variants such as activation checkpointing.
By Kunjal Panchal, Sunav Choudhary, Yuriy Brun, Hui Guan
arXiv:2607. 01125v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables fine-tuning large language models when backpropagation is unavailable or memory-prohibitive, but existing methods often perturb full model weights or randomly constructed low-dimensional subspaces, yielding high-variance estimates and limited performance.
By Xun Dong, Yibo Xu, Naigang Wang, Xin Li, Penghang Yin, Zi Yang
arXiv:2606. 17526v1 Announce Type: new Abstract: Efficient optimization is essential for training large language models.
By Da Chang, Ganzhao Yuan
The paper introduces ZFO, a lightweight framework that separates direction selection from step-size determination in large‑scale neural network optimization. ZFO uses a trusted first‑order optimizer to pick a search direction and then performs only two additional objective evaluations to build a local curvature‑aware model, selecting an adaptive step within a bounded interval. The authors provide theoretical guarantees for reliable curvature estimation, near‑optimal step selection, and convergence to a stationary point, and demonstrate that ZFO improves optimization and final performance over fixed‑step first‑order baselines on language‑model fine‑tuning tasks.
By Cristian McGee, El Houcine Bergou, Aritra Dutta
The paper introduces Ladder Side Tuning (LST), a parameter‑efficient fine‑tuning method that adds a lightweight side network to large language models. LST matches QLoRA’s compute scaling while halving peak memory usage, enabling 7B‑parameter models to be fine‑tuned on a single 12 GB GPU with 2k‑token contexts without gradient checkpointing. The authors also present xLadder, a depth‑extended variant that increases effective depth through cross‑connections, allowing deeper reasoning without extra memory overhead.
By Estelle Zheng, Nathan Cerisara, S\'ebastien Warichet, Emmanuel Helbert, Christophe Cerisara
arXiv:2502. 10239v3 Announce Type: replace-cross Abstract: Federated Learning (FL) is a promising paradigm for finetuning Large Language Models (LLMs) across distributed data sources while preserving data privacy.
By Mohamed Aboelenien Ahmed, Kilian Pfeiffer, Ramin Khalili, Heba Khdr, J\"org Henkel
The paper introduces EoupCT, a framework that estimates and orthogonalizes unknown pre‑training gradients to mitigate catastrophic forgetting during continual fine‑tuning of large language models. It generates pseudo data most susceptible to forgetting using a learnable soft prompt with Gumbel‑Softmax, then jointly optimizes model parameters and the prompt via a first‑order Pareto optimizer to enforce orthogonality between new task updates and the estimated gradients. Experiments on multiple LLMs show that EoupCT preserves both task‑specific performance and the models’ inherent general‑purpose knowledge.
By Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li, Gang Niu, Masashi Sugiyama
arXiv:2502. 11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining.
By Guoxia Wang, Shuai Li, Congliang Chen, Jinle Zeng, Jiabin Yang, Dianhai Yu, Yanjun Ma, Li Shen
arXiv:2609.38095v1 Announce Type: new
Abstract: Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU...
By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e