arXiv Machine Learning By Kejian Zhu, Zhuoran Jin, Shangqing Tu, Hongbang Yuan, Yushi Bai, Kang Liu, Juanzi Li, Jun Zhao

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

Read the original on arXiv Machine Learning →

arXiv:2608. 03573v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.