arXiv AI By Mengqi Li, Lei Zhao, Anthony Man-Cho So, Ruoyu Sun, Xiao Li

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

Read the original on arXiv AI →

arXiv:2510. 18814v4 Announce Type: replace-cross Abstract: Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training?

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.