arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
By Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf
arXiv:2608. 07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment.
By Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma
The paper introduces a two‑stage training framework for compact instruction‑following rerankers. Stage 1 strengthens a 4B teacher reranker with off‑policy GRPO using LLM‑judge feedback on 88K examples, while Stage 2 trains a 1B student by sampling its own rankings and receiving soft teacher‑derived rewards, blending exploration with knowledge transfer. The method achieves superior nDCG and MRR scores on MAIR‑11 and MAIR‑Full benchmarks, outperforming offline distillation baselines and larger RL‑trained rerankers.
By Vignesh Prabhakar, Jialing Pan, Anil Babu Ankisettipalli
Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of exampl...
arXiv:2607. 27787v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is identically zero, so GRPO produces no gradient on precisely the prompts at the frontier of the model's capability.
By Ken Ding
arXiv:2510. 14807v3 Announce Type: replace Abstract: We revisit exploration collapse in reinforcement learning with verifiable rewards (RLVR), from the perspective of the \emph{candidate distribution} for next-token prediction.
By Ruotian Peng, Yi Ren, Zhouliang Yu, Weiyang Liu, Yandong Wen