arXiv Machine Learning By Huaqing Zhang, Jingchu Gai, Juno Kim, Bingbin Liu, Andrej Risteski

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon

Read the original on arXiv Machine Learning →

arXiv:2606. 30445v1 Announce Type: new Abstract: Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-training approach, often outperforming offline supervised fine-tuning (SFT).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.