arXiv AI By Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

Read the original on arXiv AI →

Rationale-Guided Policy Optimization (RGPO) is a reinforcement‑learning framework that adaptively uses ground‑truth rationale information to scaffold a language model’s reasoning process. Instead of treating reference solutions as fixed imitation targets, RGPO temporarily incorporates rationales to help the model generate better responses, then reverts to unguided learning with higher‑reward, model‑generated solutions. Experiments in both language‑only and vision‑language tasks show that RGPO consistently outperforms RLVR baselines, with ablation studies confirming that adaptive rationale guidance is a key factor in its success.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 19

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

arXiv:2510. 21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models.

By Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei
arXiv AI
Jun 24

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

arXiv:2606. 24064v1 Announce Type: new Abstract: Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason.

By Tianyuan Shi, Canbin Huang, Bei Li, Xin Chen, Xiaojun Quan, Jingang Wang, Qifan Wang