Hugging Face Trending Papers

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

Read the original on Hugging Face Trending Papers →

Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it inherently suffers from exposure bias due to the privileged information of the self-teacher.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Jul 31

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

arXiv:2607. 28026v1 Announce Type: new Abstract: Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD).

By Xingjian Wu, Junlin Liu, Xingchen Liu, Xuhang Zhu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai
arXiv AI
4d ago

SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

SIPO (Self‑Instructing Policy Optimization) unifies reinforcement learning with on‑policy self‑distillation by using a contrastive self‑teacher to generate token‑level credit signals. The method samples multiple rollouts per prompt, pairs each with a reference answer and its mistakes, and uses the difference in teacher log‑probabilities to provide dense feedback while still respecting the overall task reward. Experiments on reasoning and code‑generation benchmarks show that SIPO outperforms both RLVR and OPSD baselines without requiring an external teacher or extra generation steps.

By Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang