arXiv:2606. 09932v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training.
By Runze Liu, Jiashun Liu, Xu Wan, Yuqian Fu, Ling Pan
TailSFT is a simple modification to supervised fine‑tuning that filters out already well‑modeled sequences, concentrating learning on the tail of the data distribution. On the OLMo‑3 7B model, this approach improves pass@16 performance on math and coding tasks by up to 17% absolute and yields up to 4% absolute gains in subsequent GRPO reinforcement‑learning runs, with only minimal computational overhead. The authors also provide a lightweight diagnostic to identify settings where TailSFT is most beneficial and argue for a stage‑aware development strategy that evaluates intermediate checkpoints by their support for later training.
By Sadhika Malladi, Samy Jelassi, Dylan Foster, Jordan T. Ash, Akshay Krishnamurthy
arXiv:2610.01133v1 Announce Type: cross
Abstract: Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yiel...
By Bangji Yang, Jiajun Fan, Hongba Ma, Ruihan Guo, Ge Liu
arXiv:2606. 15333v1 Announce Type: cross Abstract: LLM unlearning has emerged as a cost-effective alternative to full retraining for removing hazardous knowledge from pretrained models while preserving general utility.
By Zirui Pang, Chenlong Zhang, Haosheng Tan, Zhuoran Jin, Jiaheng Wei, Zixin Zhong
arXiv:2609.37169v1 Announce Type: cross
Abstract: Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since add...
By Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu, Zhiqiang Zhang, Xiaolin Huang, Jun Zhou
arXiv:2606. 03073v1 Announce Type: cross Abstract: Reinforcement learning (RL) for large language models (LLMs) is highly sensitive to hyperparameter configurations, making hyperparameter optimization (HPO) essential yet computationally expensive.
By Minping Chen, Bowen Xiao, Du Liang, Chuxuan Zeng, Zeyi Wen