The paper introduces Group‑Calibrated On‑Policy Distillation (GC‑OPD), a method that aligns token‑level teacher guidance with trajectory‑level verifier rewards for long‑context reasoning tasks. GC‑OPD normalizes rewards within rollout groups, uses the signed teacher‑verifier disagreement as a residual, and distributes this residual across tokens via Relative‑Advantage‑Based Credit Assignment (RACA). Experiments on five long‑context benchmarks show that GC‑OPD improves Qwen3‑4B and Qwen3‑8B checkpoints from 29.08/35.12 to 40.47/44.65, outperforming vanilla OPD and demonstrating the effectiveness of group‑relative residual calibration.
By Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou
arXiv:2607. 07050v3 Announce Type: replace-cross Abstract: Top-K teacher logits make on-policy distillation tractable, but probability mass is not the same as decision support.
By Jiabin Shen, Guang Chen, Chengjun Mao
arXiv:2608. 09447v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation.
By Zehao Chen, Gongxun Li, Tianxiang Ai, Yifei Li, Zixuan Huang, Wang Zhou, Tao Huang, Fuzhen Zhuang, Xianglong Liu, Jianxin Li, Deqing Wang, Yikun Ban
The study investigates how multi‑harness reinforcement learning (RL) affects coding agents by comparing two grouping strategies—Within (one group per task‑harness pair) and Cross (harnesses pooled within a task)—using a Qwen3‑8B policy trained on frozen task‑harness records from Aider, OpenHands, Qwen Code, and SWE‑agent. Across 24,000 sealed evaluations, the choice of evaluation harness dramatically increases solve rates (from 2.14 % to 9.27 %), while the grouping rule has a negligible effect. Both grouping rules yield similar gains on the same source harness, and Cross‑harness credit does not improve portability beyond Within‑harness credit, suggesting that multi‑harness RL reports should specify grouping boundaries and test on unseen harnesses.
By Chenqian Le, Jiayi Cheng, Qijia He, Runhao Li, Yinghao Li, Xupeng Chen
The paper introduces Group-Calibrated On-Policy Distillation (GC‑OPD), a method that aligns token‑level teacher guidance with task‑level verifier rewards for long‑context reasoning. GC‑OPD normalizes verifier and OPD scores within rollout groups, uses their difference as a signed disagreement residual, and redistributes this residual across tokens via Relative‑Advantage‑Based Credit Assignment (RACA). Experiments on five long‑context benchmarks show that GC‑OPD improves Qwen3‑4B and Qwen3‑8B checkpoints from 29.08/35.12 to 40.47/44.65, outperforming vanilla OPD and demonstrating the effectiveness of group‑relative residual calibration.
arXiv:2607. 21273v2 Announce Type: replace Abstract: Dense per-step supervision is the standard remedy for sparse-reward long-horizon LLM agents: reward the policy for predicting its next observation, which looks provably safe under potential-based shaping.
By Yu Wang
arXiv:2605. 31191v2 Announce Type: replace Abstract: We investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on CIFAR-10.
By Umut Onur Yasar
arXiv:2608. 09826v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.
By Yubo Jiang, Fengying Xie, Zhiguo Jiang, Haopeng Zhang
The paper introduces Selective Supervision for Direct-OPD (S$^2$D-OPD), a refinement of Direct On-Policy Distillation that filters out states where the teacher’s policy change is minimal, as measured by the teacher‑reference Jensen‑Shannon divergence. By masking low‑divergence states and keeping only the top 10% of states per response, S$^2$D-OPD improves held‑out accuracy on AIME and HMMT benchmarks across multiple teacher‑student pairs without additional forward passes.
By Yibo Zhao, Zixuan Yang, Yunshi Lan, Xiang Li
Agentic systems have widened the gap between producing candidate outputs and reviewing them. This paper asks a practical architectural question: should domain specialization be built into an evaluator's weights, or into the rule that decides when its judgment can be trusted?
arXiv:2608. 11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.
By Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
arXiv:2608.21098v1 Announce Type: new
Abstract: Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or har...
By Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva