arXiv Machine Learning By Hongyu Chen, Xinyi Luo, Ming Zhao, Lin Tang, Zihan Xu, Jing Li, Yuxuan Wang, Haoran Deng, Wei Zhang

Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning

Read the original on arXiv Machine Learning →

arXiv:2610. 00436v1 Announce Type: new Abstract: Online batch selection fine-tunes a language model on the most useful part of each candidate batch.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

A Comedy of Estimators: On KL Regularization in RL Training of LLMs

The paper investigates how different estimators of the reverse Kullback–Leibler (KL) divergence used as a regularization term in reinforcement learning (RL) training of large language models (LLMs) affect training stability and downstream performance. By analyzing gradient bias across various estimator configurations, the authors demonstrate that biased gradients can cause training instabilities, while unbiased configurations improve performance on both in‑domain and out‑of‑domain tasks. Experiments on Qwen2.5‑7B, Llama‑3.1‑8B‑Instruct, and Qwen3‑4B‑Instruct‑2507 confirm these findings and show that KL regularization also stabilizes off‑policy RL training in asynchronous setups.

By Vedant Shah, Johan Obando-Ceron, Vineet Jain, Brian Bartoldson, Bhavya Kailkhura, Sarthak Mittal, Glen Berseth, Pablo Samuel Castro, Yoshua Bengio, Esmeralda S. Whitammer, Moksh Jain, Siddarth Venkatraman, Aaron Courville