arXiv Machine Learning By Rachit Bansal, Clara Mohri, Tian Qin, David Alvarez-Melis, Sham Kakade

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

Read the original on arXiv Machine Learning →

arXiv:2606. 04272v1 Announce Type: new Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

TailSFT is a simple modification to supervised fine‑tuning that filters out already well‑modeled sequences, concentrating learning on the tail of the data distribution. On the OLMo‑3 7B model, this approach improves pass@16 performance on math and coding tasks by up to 17% absolute and yields up to 4% absolute gains in subsequent GRPO reinforcement‑learning runs, with only minimal computational overhead. The authors also provide a lightweight diagnostic to identify settings where TailSFT is most beneficial and argue for a stage‑aware development strategy that evaluates intermediate checkpoints by their support for later training.

By Sadhika Malladi, Samy Jelassi, Dylan Foster, Jordan T. Ash, Akshay Krishnamurthy