arXiv AI By Zeju Qiu, Lixin Liu, Adrian Weller, Han Shi, Weiyang Liu

POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation

Read the original on arXiv AI →

arXiv:2603. 05500v2 Announce Type: replace-cross Abstract: Efficient and stable training of large language models (LLMs) remains a core challenge in modern machine learning systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 27

Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets

The paper introduces Ladder Side Tuning (LST), a parameter‑efficient fine‑tuning method that adds a lightweight side network to large language models. LST matches QLoRA’s compute scaling while halving peak memory usage, enabling 7B‑parameter models to be fine‑tuned on a single 12 GB GPU with 2k‑token contexts without gradient checkpointing. The authors also present xLadder, a depth‑extended variant that increases effective depth through cross‑connections, allowing deeper reasoning without extra memory overhead.

By Estelle Zheng, Nathan Cerisara, S\'ebastien Warichet, Emmanuel Helbert, Christophe Cerisara